יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SimpleMemVLA: זיכרון פשוט ומוצלח למודלי תצלום-לשון-פעולה

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
SimpleMemVLA הוא זיכרון חדשני למודלי תצלום-לשון-פעולה. הוא משתמש במצב התצלום המקורי כזיכרון, ומצליח להשיג תוצאות טובות בביצועים.
תקציר מקורי באנגליתarXiv:2609.05533v2 Announce Type: replace-cross Abstract: Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, must decide what to keep from the past before knowing what a future decision will require. They were motivated by the assumption that minute-scale history is too large to process directly, which no longer holds for modern VLM backbones. We propose SimpleMemVLA, a VLA without a dedicated memory module that uses the backbone's native video context directly as memory. It keeps the sampled history intact in the timestamped video format the backbone was pretrained to process, routes t
קרא במקור המקורי