כתבה
arXiv cs.LG ·
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
תקציר מקורי באנגליתarXiv:2609.36812v1 Announce Type: new Abstract: Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions un
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית