כתבה
arXiv cs.CL ·
למידה מיתרון מצב חשיבה
Learning from Think-Mode Advantage via On-Policy Distillation
חוקרים פיתחו שיטה חדשה ללמידה מיתרון מצב חשיבה במודלים גדולים של שפה. השיטה, הנקראת On-Policy Distillation, מאפשרת למודלים ללמוד מניסיון קודם ולשפר את ביצועיהם. החוקרים בדקו את השיטה במספר מודלים, כולל LLaMA, ומצאו שהיא משפרת את הביצועים במגוון משימות.
תקציר מקורי באנגליתarXiv:2609.37044v2 Announce Type: replace Abstract: Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית