כתבה
arXiv cs.AI ·
SIPO: יחדוע רכיבת השלמה עצמית עם פענוח תגמול
SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
SIPO: יחדוע רכיבת השלמה עצמית עם פענוח תגמול. המחקר מציע פתרון לבעיית חירות תגמול על ידי שימוש במודל עצמי. המודל העצמי נותן תגמול עדין למודל המקורי, מאפשר לו ללמוד ולשפר. המחקר מדגים את יעילות הפתרון במספר תחומים, כולל רכיבת השלמה עצמית ופענוח תגמול.
תקציר מקורי באנגליתarXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית