כתבה
arXiv cs.AI ·
איחוד מודלים: שיקום נגדי בעיצוב מדיניות
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
חוקרים פיתחו שיטה חדשה לשיפור ביצועי מודלים באמצעות עיצוב מדיניות עם שיקום נגדי. השיטה מאפשרת למודלים ללמוד מטעויות ולשפר את ביצועיהם. החוקרים בדקו את השיטה במספר סימולציות וגילו שהיא משפרת משמעותית את הביצועים.
תקציר מקורי באנגליתarXiv:2608.04408v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conventionally supervises the corresponding trajectory. On AIME branch diagnostics, the mean continuation-minus-rollback effect is 0.185 for recoverable states and -1.000 for irreversible-but-avoidable states, demonstrating opposite intervention
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית