כתבה
arXiv cs.LG ·
הפחתת המיסוי של הגדילה באורך
Mitigating the Length-Scaling Tax with Online Distillation
אנו הציגנו טכניקה להפחתת המיסוי של הגדילה באורך, כדי לשמור על תגובות קצרות על שאלות פשוטות.
תקציר מקורי באנגליתarXiv:2609.38854v1 Announce Type: new Abstract: Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית