יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Looped GPT-BERT: סחר בפרמטרים למחשב בקטעי שפה קטנים

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling
כשמידע האוריינטלי מוגבל, גדילת פרמטרים אינה הדרך היחידה לשפר את יכולת הדגימה של דגם שפה. ניתן לשפר את הדגימה על ידי חזרה על פרמטרים קטנים. ניתן להשתמש ב-GPT-BERT כדי לשפר את יכולת הדגימה של דגם שפה.
תקציר מקורי באנגליתarXiv:2609.09691v1 Announce Type: new Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare objective ratios, non-looped and looped architectures, and loop counts. Our final $4\times12$ model uses four physical layers for twelve recurrent traversals and contains 12.18M parameters. The BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. Compared with pub
קרא במקור המקורי