כתבה
arXiv cs.LG ·
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
הכנה חום ללמידת ריפוד בעזרת דיסטילציה של פוליציה
תקציר מקורי באנגליתarXiv:2610.02781v1 Announce Type: new Abstract: Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the sec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית