כתבה
arXiv cs.LG ·
אינטרפולציה אוניברסלית לרשתות עצביות עמוקות
Universal interpolation for deep residual self-attention networks
חוקרים הוכיחו אינטרפולציה אוניברסלית לרשתות עצביות עמוקות עם תשומות קצרות. התוצאה מתקבלת עם שני בלוקים קפואים בלבד, ומאפשרת העברה בין רצפים שונים.
תקציר מקורי באנגליתarXiv:2610.01981v1 Announce Type: new Abstract: Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית