כתבה
arXiv cs.AI ·
DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance
DOHF מציע גידונג אונליין של תהליך דיפוזיה לאימון תוך-זמן, עם התאמה למדלי Doob's $h$-transform. זה מאפשר אימון תוך-זמן של גנרטיבי עם תכונות רגולטוריות.
תקציר מקורי באנגליתarXiv:2609.31882v2 Announce Type: replace-cross Abstract: Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-different
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית