יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

RLCSD: שיפור בלמידת רפלקסיה עם עקרון התנגדות

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
מחקר חדש: RLCSD - שיפור בלמידת רפלקסיה עם עקרון התנגדות. פיתוח חדשני למודלי תפיסה, המשפר את הביצועים שלהם. המחקר נעשה על ידי צוות מדענים בארגון Qwen, והוא כולל שימוש במודלי GPT-5.
תקציר מקורי באנגליתarXiv:2606.11709v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified solution. However, we show that the resulting distributional gap concentrates on style tokens rather than task-bearing ones, as the hinted model tends to produce shorter, more direct outputs. We term this pathology \emph{privilege-induced style drift}, which can destabilize training and shorten responses. To address this, we propose \textbf{RLCSD} (Reinforcement Learning with Contrastive on-policy Self-Distillation), which mitigates this drift by contrasting the teacher-student gap under a correct hint against that under a wrong hint, suppressing sty
קרא במקור המקורי