כתבה
arXiv cs.LG ·
TV-Regulated OPD: Direction Matters in On-Policy Distillation
תקציר מקורי באנגליתarXiv:2609.08341v1 Announce Type: new Abstract: On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית