כתבה
arXiv cs.LG ·
Q-Learning with Scalar Adjoint Matching
תקציר מקורי באנגליתarXiv:2610.10437v1 Announce Type: new Abstract: Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית