יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

טרנספורמר מחזורי

Recurrent Looped Transformer
המחברים מציגים את ה-Recurrent Looped Transformer (RLT), מודל שמשלב מרכיבים מקבילים ומחזוריים. RLT מראה יכולת טובה יותר במשימות עקביות ארוכות, כגון עקביות פריטים וחישובים מודולריים.
תקציר מקורי באנגליתarXiv:2610.07591v1 Announce Type: new Abstract: State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based $S_5$ permutation track
קרא במקור המקורי