כתבה
arXiv cs.CL ·
RLCSD: רכיבת למידת רפלקסיה עם קונטרסטיביות על-מדרגת עצמ-הפרדה
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
RLCSD: רכיבת למידת רפלקסיה עם קונטרסטיביות על-מדרגת עצמ-הפרדה. המחקר פותח טכניקה חדשה ללמידת רפלקסיה שמצליחה למנוע דחיפה של סגנון עקב הסימנים. הטכניקה, RLCSD, נבחנה במודלי Qwen3 ו- Olmo-3-7B-Think והציגה תוצאות טובות יותר משיטות קודמות.
תקציר מקורי באנגליתarXiv:2606.11709v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified solution. However, we show that the resulting distributional gap concentrates on style tokens rather than task-bearing ones, as the hinted model tends to produce shorter, more direct outputs. We term this pathology \emph{privilege-induced style drift}, which can destabilize training and shorten responses. To address this, we propose \textbf{RLCSD} (Reinforcement Learning with Contrastive on-policy Self-Distillation), which mitigates this drift by contrasting the teacher-student gap under a correct hint against that under a wrong hint, suppressi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית