כתבה
arXiv cs.CL ·
SIPO: חברת הלמידה המשותפת: חיזוק שיטת הלמידה על ידי עצמאיות עצמית
SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
SIPO: שיטת הלמידה המשותפת, שמשלבת חיזוק שיטת הלמידה עם עצמאיות עצמית, נועדה לפצל קרדיט על פי תווים. השיטה, המכונה SIPO, משתמשת במודל עצמאי כדי לספק סימני למידה צפוף. במחקר זה, נבחנה SIPO במספר בסיסים של תפיסה והפקת קוד, והתברר כי SIPO משפרת את התוצאות יותר משיטות הלמידה המקוריות.
תקציר מקורי באנגליתarXiv:2609.36742v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית