כתבה
arXiv cs.LG ·
Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention
תקציר מקורי באנגליתarXiv:2609.39494v1 Announce Type: cross Abstract: Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית