כתבה
arXiv cs.AI ·
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2609.39820v1 Announce Type: cross Abstract: Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into correctiv
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית