כתבה
arXiv cs.LG ·
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
תקציר מקורי באנגליתarXiv:2605.15055v2 Announce Type: replace Abstract: Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization. Extending RL to multiple tasks is challenging: joint optimization suffers from cross-task interference and imbalance, while cascade RL is cumbersome and prone to catastrophic forgetting. We propose DiffusionOPD, a new multi-task training paradigm for diffusion models based on Online Policy Distillation (OPD). DiffusionOPD first trains task-specific teachers independently, then distills their capabilities into a unified student along the student own rollout trajectories. This decouples single-task exploration from multi-task integration and avoids the optimization bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית