כתבה
arXiv cs.LG ·
חשיפת תהליך הצפייה: כאשר הוא עוזר, כאשר הוא פוגע ולמה
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
תהליך הצפייה: כאשר הוא עוזר, כאשר הוא פוגע ולמה. ניתוח של תהליך הצפייה, כאשר הוא עוזר, כאשר הוא פוגע ולמה. ניתוח של תהליך הצפייה, כאשר הוא עוזר, כאשר הוא פוגע ולמה.
תקציר מקורי באנגליתarXiv:2605.10889v2 Announce Type: replace Abstract: On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically requires costly training runs whose aggregate performance metrics obscure the dynamics at the level of individual tokens. We introduce a training-free diagnostic framework that operates at the highest resolution: per token, per question, and per teacher. We derive an ideal per-node gradient defined as
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית