יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הדיסטילציה בעקבות תגמול ניתן לאימות

On-policy Distillation with Verifiable Reward
הצגת תגמול ניתן לאימות והדיסטילציה בעקבות תגמול כדי לשפר את ביצועי המודלים לשפה גדולים לאחר הכשרה.
תקציר מקורי באנגליתarXiv:2608.24696v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OP
קרא במקור המקורי