יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ניווט ראייה-שפה עם תנאי מצב עתידי

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
ניווט ראייה-שפה עם תנאי מצב עתידי (FSC-VLN) משפר את היכולת לנווט בסביבה. המודל משתמש באותות עתידיים כדי לשפר את הניווט. FSC-VLN מראה שיפור בביצועים על רקע StreamVLN.
תקציר מקורי באנגליתarXiv:2607.18042v2 Announce Type: replace-cross Abstract: End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly encourage the policy state to be predictive of future visual outcomes, limiting long-horizon decision making. A privileged-input diagnostic shows that access to an expert-trajectory future image can substantially improve navigation, indicating that future observations contain rich, actionable cues, though such inputs are unavailable at deployment. Motivated by this signal, we propose Future-State-Conditioned VLN (FSC-VLN), a deployable model that augments a causal policy with a future-query token and uses
קרא במקור המקורי