כתבה
arXiv cs.CL ·
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
תקציר מקורי באנגליתarXiv:2609.39827v1 Announce Type: new Abstract: Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית