יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

תקציר מקורי באנגליתarXiv:2610.02781v1 Announce Type: cross Abstract: Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the s
קרא במקור המקורי