כתבה
arXiv cs.AI ·
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
תקציר מקורי באנגליתarXiv:2610.01789v1 Announce Type: cross Abstract: Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית