יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

MIKASA-Robo-VLA: בנצ'מרק לזיכרון במודלים VLA

MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
MIKASA-Robo-VLA הוא בנצ'מרק לבדיקת זיכרון במודלים VLA. הוא כולל 90 משימות שונות, רובן מחייבות זיכרון. הבנצ'מרק מורכב מ-22,500 נתיבים אורקוליים. נבדקת גם גרסת בסיס עם תמונות נוכחיות ותנועה, אך ללא היסטוריית תצפית או מודול זיכרון מפורש.
תקציר מקורי באנגליתarXiv:2610.00604v1 Announce Type: new Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap cou
קרא במקור המקורי