יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור מדיניות דיפוזיה

Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
שיטה חדשה לשיפור מדיניות דיפוזיה בלמידת חיזוק. השיטה משלבת בחירת הצעות עם זרימת עדכון מותנית. תוצאות מבטיחות ב-50 משימות OGBench.
תקציר מקורי באנגליתarXiv:2609.36812v2 Announce Type: replace-cross Abstract: Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final
קרא במקור המקורי