יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

על הפיכת סקאלה של טרנספורמרים סילוניים: יציבות ועבריינות

On the Residual Scaling of Looped Transformers: Stability and Transferability
במאמר זה, המחברים חוקרים את הפיכת סקאלה של טרנספורמרים סילוניים ומציעים פתרון לבעיית היציבות והעבריינות שלהם. הם מציעים פתרון חדש לבעיית היציבות של טרנספורמרים סילוניים, שמאפשר עבריינות טובה יותר וקצב קלט-פלט גבוה יותר.
תקציר מקורי באנגליתarXiv:2606.18524v2 Announce Type: replace Abstract: Looped (weight-tied) Transformers apply a shared residual block $N$ times ($h \leftarrow h + \varepsilon\,f(h)$, same $f$ at each step), increasing effective depth without adding parameters. Prior depth-scaling analyses prescribe $\varepsilon = 1/\!\sqrt{L}$ for depth-$L$ residual networks. We show that this is insufficient for looped architectures: weight sharing makes residual updates correlated across iterations, requiring the stronger scaling $\varepsilon = 1/N$. For multi-layer blocks ($L$ unique layers looped $N$ times), we derive a factored parameterization $\varepsilon = \lambda/(N\!\sqrt{L})$ that separates the two sources of growth: $1/N$ controls the within-layer loop correlation, and $1/\!\sqrt{L}$ controls the across-layer va
קרא במקור המקורי