כתבה
arXiv cs.AI ·
Learned Image Compression for Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2606.16253v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in bandwidth-constrained or distributed deployment settings. Existing image and video codecs, however, are designed to preserve generic visual fidelity rather than the control performance of downstream VLA policies. In this work, we introduce SPARC (SPatially Adaptive Rate Control), a learned image compression framework tailored for VLA-driven robots. Our key observation is that the importance of visual information varies substantially across both camera views and spatial regions within an image. Based on this observation, SPARC employs a lightweight temporal
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית