יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Agentic-TTT: טיפול במדיניות זמן המבחן לטיפול בזמן המבחן

Agentic-TTT: Training test-time policy for test-time training
Agentic-TTT לומדת מדיניות זמן המבחן לשלוט בהחלטות טיפול בזמן המבחן. המדיניות נלמדת על ידי הצפייה בתועלת הצפויה מהחלטותיה. Agentic-TTT מכפילה את התועלת על פני המודל הבסיסי, ולומדת להסתכל בין תועלת למחשוב.
תקציר מקורי באנגליתarXiv:2610.12002v1 Announce Type: cross Abstract: Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which le
קרא במקור המקורי