כתבה
arXiv cs.AI ·
Guided Action Flow: Value-Guided Sampling for Frozen Vision-Language-Action Policies
תקציר מקורי באנגליתarXiv:2607.02092v4 Announce Type: replace-cross Abstract: Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves further updates to the policy parameters. For flow-matching policies, iterative action generation provides an additional opportunity to incorporate task information during inference. We introduce Guided Action Flow (GAF), which learns a compact, observation-conditioned action-value critic from robot task rollouts and applies its action gradient to steer reverse-time flow sampling. The supervised-fine-tuned VLA remains frozen throughout critic learning and deployment. Physical-robot experiments show an increase in aggregate success from 60.0% to 82.5% across six nominal manipulation tasks. Under six
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית