כתבה
arXiv cs.AI ·
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
תקציר מקורי באנגליתarXiv:2605.21862v3 Announce Type: replace-cross Abstract: Chunked vision-language-action (VLA) policies generate multi-step actions from one observation and typically re-observe after executing the action chunk. Earlier actions change the scene conditions for later steps, while occlusion can compromise visual feedback. Spatial and temporal VLAs enhance geometry and memory, but their representations may become stale as actions alter object poses and contacts. We introduce EvoScene-VLA, which uses compact scene tokens to unify within-chunk scene prediction, cross-chunk state propagation, and observation-based correction. The action expert jointly denoises actions and future scene states in a single flow-matching process, allowing predicted scene changes to inform action generation. The final
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית