יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

איזון חד-צדדי לאסטימציה Top-$k$ להטמעה על-מדיני

Unbiased Top-$k$ Estimation for On-Policy Distillation
איזון חד-צדדי לאסטימציה Top-$k$ להטמעה על-מדיני. המחברים מציגים פתרון לאי-איזון של אסטימציה Top-$k$ בהטמעה על-מדיני.
תקציר מקורי באנגליתarXiv:2609.34447v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent wor
קרא במקור המקורי