כתבה
arXiv cs.LG ·
מספרי ספקטרום לסדרי זמן משותפים באימון LLM: 3+3(+2) תהליכי סקאלה
From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes
חוקי סקאלה באימון LLM קדם-למודל ב-SGD עייפה-אונליין רעשי
תקציר מקורי באנגליתarXiv:2609.40148v1 Announce Type: new Abstract: Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time $T_t=\sum_{s<t}\eta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית