כתבה
arXiv cs.AI ·
GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales
תקציר מקורי באנגליתarXiv:2609.39692v1 Announce Type: cross Abstract: On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית