כתבה
arXiv cs.LG ·
הנהיגה של מודלי ריצה חוזרת בזמן המודליזציה עם תגובה של קריאה
Steering Recurrent Reasoners at Inference Time with Readout Feedback
מודלי ריצה חוזרת יכולים להיות משופרים בזמן המודליזציה עם תגובה של קריאה. ניתן להשתמש בסיכויי קריאה של המודל כדי להנהיג את הדינמיקה הלוואי. תוצאות טובות נרשמו במספר ניסויים.
תקציר מקורי באנגליתarXiv:2608.24136v2 Announce Type: replace Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more steps or sampling more trajectories, but ignore information revealed within each trajectory. Here we show that recurrent models can be improved at inference time by using their own readout probabilities to steer latent dynamics without retraining. We introduce Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics. Across three recurrent models (AKOrN, ItrSA++, TRM) on Sudoku and Maze, RoFB yields clear ga
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית