כתבה
arXiv cs.LG ·
Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2610.10526v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $\pi_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $\pi_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is sy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית