כתבה
arXiv cs.LG ·
Learning from Hindsight for VLA Reinforcement Learning
תקציר מקורי באנגליתarXiv:2607.09042v2 Announce Type: replace Abstract: Reinforcement learning is increasingly used to fine-tune vision-language-action (VLA) models, but robot interaction is expensive and learning becomes highly sample inefficient when successful rollouts are rare. When reward is assigned only for completing the commanded task, a failed rollout is treated as having no value even if it successfully executes behaviors relevant to that task. A robot that fails to place the correct object in a bowl may still move that object toward the bowl or place a different object inside it, demonstrating objects and actions that can be reused to solve the target task. These behaviors define auxiliary tasks that the policy can already solve, providing useful learning signals even before it can solve the harde
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית