כתבה
arXiv cs.CL ·
למידה מטעויות שמשתנות: תיקון אדפטיבי איטרטיבי להתפשטות על-מדיני
Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
AIR-OPD - תיקון אדפטיבי איטרטיבי להתפשטות על-מדיני - משפר את החשיבה המתמטית. המאמר עוסק באימון מודלי LLM לצורך תיקון שגיאות. המודלים Claude ו-GPT-5 הוצגו כמודלים המשמשים לצורך תיקון. המאמר עוסק בשיפור הביצועים של המודלים.
תקציר מקורי באנגליתarXiv:2610.02700v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guida
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית