כתבה
arXiv cs.CL ·
למידה מתוך תיקון תקיפה: Root-Cause-Guided On-Policy Distillation
Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation
RC-OPD מציע שיטה חדשה לשיפור תקיפת המוח של ה-AI. השיטה משתמשת בתיקונים של תקיפת המוח של התלמיד כדי לספק הנחיות שמתייחסות לשגיאותיו. RC-OPD מציע שיפור ניכר בביצועי ה-AI.
תקציר מקורי באנגליתarXiv:2610.03515v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית