כתבה
arXiv cs.AI ·
SR-TTT: תיקון וניתוח מנגנוני
SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
תיקון לSR-TTT: הוכח שהשיפורים המדווחים היו תוצאה של בעיות במדדים. המחברים מציעים מימוש מתוקן ומנתחים את הכישלון לשני מכשולים עצמאיים.
תקציר מקורי באנגליתarXiv:2603.06642v2 Announce Type: replace-cross Abstract: Test-Time Training (TTT) language models replace the KV-cache with fast weights updated during inference, achieving O(1) memory but suffering catastrophic failure on exact-recall tasks. Version 1 of this work proposed SR-TTT, which routes high-surprisal tokens to a sparse exact-attention Residual Cache, and reported large Needle-in-a-Haystack gains. We show those gains were evaluation artifacts: the loss and metric read logits at the answer positions rather than one position earlier, training both models to copy an answer already visible in their input (a model trained on retrieval-impossible data reaches 100% accuracy under the flawed metric); additionally, the cache attended non-causally over future tokens, including the answer it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית