כתבה
arXiv cs.AI ·
Show-Harness: רק גוף ולמידה ולמידה יכולים לשחק רובוטים
Show-Harness: Just a VLM Agent Can Play Robots
מודלי VLM הבסיסיים יכולים לשלוט על רובוטים דרך ממשק סמנטי קצר. Show-Harness מציג פתרון זה, המאפשר ל-VLMs ל"שחק" רובוטים דרך ממשק סמנטי קצר. הממשק מקשר בין כוונה לפעולה, והוא ניתן להסברה ולהסברה. Show-Harness מדגים את האפשרות להפעיל VLMs סגור-מקור לשליטה על רובוטים באופן זרוע-שטח, ולהתאים VLMs קטני-ממד להתקנה זולה עם רק כמה GPU-שעות של התאמה.
תקציר מקורי באנגליתarXiv:2609.10522v1 Announce Type: cross Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-sou
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית