כתבה
arXiv cs.LG ·
SimpleMemVLA: זיכרון פשוט ומוצלח למערכות תצוגה-שפה-פעולה
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
נוסחא חדשה לזיכרון למערכות תצוגה-שפה-פעולה, SimpleMemVLA, שמשמשת כזיכרון פשוט ומוצלח. המאמר עוסק בפיתוח זיכרון חדש למערכות VLA, שמספק זיכרון פשוט ומוצלח למערכות VLA. הזיכרון החדש, SimpleMemVLA, מספק זיכרון פשוט ומוצלח למערכות VLA, ומשמש כזיכרון פשוט ומוצלח למערכות VLA.
תקציר מקורי באנגליתarXiv:2609.05533v1 Announce Type: cross Abstract: Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms: retrieval banks, learned compressors, recurrent states must decide what to keep from the past before knowing what a future decision will require. This was motivated by the assumption that minute-scale history is too large to process directly, which modern VLM backbones no longer make true. In this work, we introduce SimpleMemVLA, a VLA without a dedicated memory module. It keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process; the hidden states of a generated sub-task then form the o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית