יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

HIL-UMI: פוסט-אימון של דגמי תצלום-שפה-פעולה באמצעות רשת יד-בכוח

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface
HIL-UMI היא רשת יד-בכוח שמאפשרת פוסט-אימון של דגמי תצלום-שפה-פעולה באמצעות רשת יד-בכוח. הרשת משתמשת באסקור תרמי כדי לזהות אזורי פחד-הפצרה ולאסוף נתונים חדשים. HIL-UMI מציע דרך יעילה יותר לפוסט-אימון של דגמי VLA.
תקציר מקורי באנגליתarXiv:2609.20659v2 Announce Type: replace-cross Abstract: Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-trai
קרא במקור המקורי