כתבה
arXiv cs.CL ·
On-Policy Distillation Teaches New Skills but Not New Knowledge
תקציר מקורי באנגליתarXiv:2610.09639v1 Announce Type: new Abstract: On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the exe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית