יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

LM-X: תיאור ניתן להבנה של רכיבי ראייה-שפה-פעולה על ידי תחזית, אירוע ואי-וודאות

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
LM-X לומד לתאר את התקדמות המשימה, את האירוע הבא ואת רמת הוודאות של הפעולה. המודל מציע תיאור ניתן להבנה של רכיבי ראייה-שפה-פעולה.
תקציר מקורי באנגליתarXiv:2608.25757v4 Announce Type: replace-cross Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanatory state is not. They provide no native account of three explanatory signals: task progress, the next semantic transition, or local command reliability. Prior work shows that progress and event structure aid long-horizon control and that uncertainty supports monitoring; however, such capabilities are typically added or extracted only after action pretraining. The field therefore lacks a VLA foundation model whose explanatory state is jointly pretrained with control. Drawing on biological sensorimotor organization, in which outcome-sensitive, event-segmente
קרא במקור המקורי