כתבה
arXiv cs.LG ·
LM-X: תיאור ניתן להבנה של תכנון ראייה-שפה-פעולה באמצעות תחזית, אירוע וביטול ספק
LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
LM-X היא תכנות ראייה-שפה-פעולה הסבירה שמציעה תחזית, אירוע וביטול ספק. היא לומדת שלושה סיגנלים ישירים: הערכות RTG, ETG ואי-וריאנס. LM-X מציעה 74.1% הצלחה ב-50 תפריטים קשים של RoboTwin2.0 ו-73.5% בשבעה תפריטים של רובוטים אמיתיים.
תקציר מקורי באנגליתarXiv:2608.25757v5 Announce Type: replace-cross Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanatory state is not. They provide no native account of three explanatory signals: task progress, the next semantic transition, or local command reliability. Prior work shows that progress and event structure aid long-horizon control and that uncertainty supports monitoring; however, such capabilities are typically added or extracted only after action pretraining. The field therefore lacks a VLA foundation model whose explanatory state is jointly pretrained with control. Drawing on biological sensorimotor organization, in which outcome-sensitive, event-segmente
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית