יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למידה מהתועלת של חשיבה-מוד-אדוונטג' דרך דיסטילציה-OPD

Learning from Think-Mode Advantage via On-Policy Distillation
במאמר זה, החוקרים חקרו את למידת התועלת של חשיבה-מוד-אדוונטג' דרך דיסטילציה-OPD. הם פיתחו שיטה חדשה, ThinkOPD, שמאפשרת למודלים ללמוד מהתועלת של חשיבה-מוד-אדוונטג' באופן יעיל. המאמר כולל תיאור של השיטה, תוצאות הניסויים והסיכויים ליישום בעתיד.
תקציר מקורי באנגליתarXiv:2609.37044v1 Announce Type: new Abstract: Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when respo
קרא במקור המקורי