כתבה
arXiv cs.LG ·
QF3: Fast Flow RL with Filtered Q-Gradients
תקציר מקורי באנגליתarXiv:2610.08789v1 Announce Type: cross Abstract: Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית