כתבה
arXiv cs.LG ·
Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
תקציר מקורי באנגליתarXiv:2610.02381v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית