יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כשהמסגרות האבדות את האות: בחינה סיבתית של תיקון ב-LLM Agents

When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
במאמר זה נבחן את תיקון השגיאות ב-LLM Agents. החידוש הוא בבחינה סיבתית של תיקון, שמאפשרת לזהות את הסיבות לשגיאות ולתקן אותן באופן יעיל. המחברים הציגו כלי חדש, CIR, שמסוגל להחלטות סיבתיות ולתקן שגיאות באופן יעיל. התוצאות הראו ש-CIR מגדיל את ההצלחה של ה-LLM Agents ב-3%.
תקציר מקורי באנגליתarXiv:2610.00372v1 Announce Type: new Abstract: Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWor
קרא במקור המקורי