כתבה
arXiv cs.LG ·
dFlowGRPO: אופטימיזציה של מדיניות למודלים דיסקרטיים
dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models
dFlowGRPO הוא כלי לאופטימיזציה של מדיניות למודלים דיסקרטיים. הוא מאפשר שימוש במגוון רחב של נתיבי הסתברות והתפלגויות מקור. המחקר מראה כי dFlowGRPO משפר את הביצועים על משימות יצירת תמונות והבנת רב-מודאלית.
תקציר מקורי באנגליתarXiv:2605.09291v2 Announce Type: replace Abstract: Discrete flow models (DFMs) are a class of flexible generative models for generating discrete data, and diffusion large language models (dLLMs) can be viewed as a special case with a specific choice of a mixture path and a masked source distribution. While several recent works have explored reinforcement learning for dLLMs, its application to more general discrete flow models remains underexplored. In this work, we present discrete Flow-GRPO (dFlowGRPO), a unified reinforcement learning framework for discrete flow models that supports a broad family of probability paths and non-masked source distributions. We derive the full trajectory probability for DFMs and formulate the denoising process as a Markov decision process, enabling dFlowGRP
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית