יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Sequential Beats Joint: על ההתנגשות בין תפיסת נתונים ו-RVLR

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
במאמר זה, המחברים מציגים שיטה חדשה לשילוב שתי שיטות עיקריות לשיפור תפיסת נתונים של LLMs: תפיסת נתונים על-מדינה ו-RVLR. השיטה, הנקראת OPD-then-RL, מציגה תפיסת נתונים על-מדינה כדי להרחיב את תפיסת הנתונים של התלמיד, ואז RLVR כדי לחדד את התפיסה. המחברים מציגים תוצאות ניסויים המציגות את יעילותה של השיטה.
תקציר מקורי באנגליתarXiv:2609.04108v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, y
קרא במקור המקורי