יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

FineART: קבוצת נתונים מסוגלת לאנוטציה ומודל תפעול-שפה-פעולה לביצוע משימות ידיים זוגיות

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
קבוצת נתונים חדשה, FineART, כוללת 40,543 תצוגות (1,718 שעות) ו-533,913 תת-משימות. המודל FineART-VLA, חוקר את התפעול-שפה-פעולה שלו ומגדיל את ההצלחה בביצוע משימות ידיים זוגיות.
תקציר מקורי באנגליתarXiv:2609.36416v2 Announce Type: replace-cross Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets with subtask labels annotate only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-trainin
קרא במקור המקורי