יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

PivotOPD: למידה להתאושש מטעויות פיקטיביות באג'נטים של תגובה

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
PivotOPD הוא פרקטיקה של חינוך על-פי תצפיות שמטרתה ללמד אג'נטים להתאושש מטעויות פיקטיביות. הפרקטיקה נבחנה באמצעות שלושה דגמי Qwen3 (8B ל-235B) והתגלה שיותר מחצי מהפסילות כללו טעות פיקטיבית. PivotOPD נועד למנוע טעויות פיקטיביות וללמד את האג'נט להתאושש מהן.
תקציר מקורי באנגליתarXiv:2609.40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that
קרא במקור המקורי