כתבה
arXiv cs.LG ·
From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
תקציר מקורי באנגליתarXiv:2610.02179v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית