כתבה
arXiv cs.AI ·
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
תקציר מקורי באנגליתarXiv:2512.02019v4 Announce Type: replace-cross Abstract: Diffusion models provide an expressive framework for sampling from complex, unnormalized distributions. In this work, we extend Maximum Entropy Reinforcement Learning (ME-RL) to diffusion-based policies by introducing Diffusion-Augmented Markov Decision Processes (DA-MDPs). DA-MDPs interpret each reverse-diffusion transition as an individual reinforcement-learning decision, while only the final denoised action is executed in the environment. Our DA-MDPs follow from a principled derivation based on the variational-inference formulation of ME-RL. By augmenting policy and target trajectories with intermediate diffusion variables, we obtain a tractable reverse-KL upper bound via the data-processing inequality. This bound decomposes acro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית