יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אחד, שני כלים: פיתוח זיכרון-נתונים בתחום התחזית האוטומטית

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction
במאמר זה, המחברים חוקרים את היחס בין כמות הנתונים לכמות הזיכרון בתחום התחזית האוטומטית. הם מציגים תאוריה חדשה שמסבירה את התופעה ומציעה דרכים לשפר את התחזיות.
תקציר מקורי באנגליתarXiv:2609.13500v1 Announce Type: cross Abstract: How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing $\mu$ for the resulting energy spectrum, we prove the minimax law $\mathfrak R^*_{\rm value}(n,B)\asymp_R \Phi_\mu(n^{-1})+\Phi_\mu(\tau_B), \Phi_\mu(t)=\int\min\{x,t\}\,\mu(\mathrm dx),$ for $n$ prediction blocks and a learned state with at most $2^B$ values. Data set the resolution $1/n$; memory sets the level $\tau_B$ reached by optimal bit allocation. The complete curve also recovers the positive spectrum. Energy-dimension pairing is essential
קרא במקור המקורי