כתבה
arXiv cs.CL ·
עומק וקנה מידה בתת-150M: JugnuLM-53M נגד JugnuLM-110M
Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
JugnuLM-110M משפר את GPT-X2-125M עם 12% פחות פרמטרים. המודל הגדול יותר משיג שיפורים בBLiMP וARC-Easy. השיפור מיוחס לקיבולת ועומק, לא לנתונים נוספים.
תקציר מקורי באנגליתarXiv:2609.14715v1 Announce Type: cross Abstract: We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query attention, RoPE, SwiGLU, RMSNorm, QK-Norm, and a z-loss; FineWeb-Edu data) and changing only the geometry to a deep-and-thin 23-layer x 576-hidden design. The larger model improves across the board -- BLiMP 78.1 -> 81.3, ARC-Easy 51.4 -> 52.5, WikiText-2 byte-perplexity 2.04 -> 1.95 -- and its 81.3% BLiMP essentially matches GPT-X2-125M (81.28) at about 12% fewer parameters. Notably the 110M model achieves this on fewer training tokens (about 8B vs 12B), so the gain is attributable to capacity and depth, not more data. Both models are deliberately conventional; this report is a clean scaling
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית