כתבה
arXiv cs.AI ·
אחד ספקטרום, שני משאבים: פיתוח זיכרון-נתונים בתחושיית חזרה
One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction
במאמר זה, נראה כי יש קשר בין נתונים וזיכרון. המחברים חקרו את הקשר בין שני המשאבים וגילו כי יש ספקטרום אחד שמשפיע על שניהם.
תקציר מקורי באנגליתarXiv:2609.13500v1 Announce Type: cross Abstract: How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing $\mu$ for the resulting energy spectrum, we prove the minimax law $\mathfrak R^*_{\rm value}(n,B)\asymp_R \Phi_\mu(n^{-1})+\Phi_\mu(\tau_B), \Phi_\mu(t)=\int\min\{x,t\}\,\mu(\mathrm dx),$ for $n$ prediction blocks and a learned state with at most $2^B$ values. Data set the resolution $1/n$; memory sets the level $\tau_B$ reached by optimal bit allocation. The complete curve also recovers the positive spectrum. Energy-dimension pairing is essential
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית