יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Puro-2B: פורו-2B: תזוזה נמוכה בעלות לאימון מודלי שפה

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
אימון מודלי שפה בעלות נמוכה באמצעות GPU צרכני. פורו-2B הוא תזוזה נמוכה בעלות לאימון מודלי שפה. הפרויקט מציע תזוזה נמוכה בעלות לאימון מודלי שפה, כולל תזוזה נמוכה בעלות לאימון מודלי שפה.
תקציר מקורי באנגליתarXiv:2608.27370v2 Announce Type: replace Abstract: Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection di
קרא במקור המקורי