כתבה
arXiv cs.LG ·
Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does
תקציר מקורי באנגליתarXiv:2609.39892v1 Announce Type: new Abstract: Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations? We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph st
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית