כתבה
arXiv cs.LG ·
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
תקציר מקורי באנגליתarXiv:2609.15883v1 Announce Type: new Abstract: Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית