יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LEGO Co-builder: מודלים רב-מודאליים לעזרי הרכבה

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
LEGO Co-builder הוא בנך מודלים רב-מודאליים לעזרי הרכבה. המחקר בוחן את היכולת של מודלים כמו GPT-4o, Gemini ו-Qwen-VL להבין הוראות הרכבה רב-מודאליות. התוצאות מראות כי גילוי אובייקטים הושג ברמה גבוהה, אך הבנת סצנה עדינה וזיהוי מצב הרכבה נותרו אתגריים.
תקציר מקורי באנגליתarXiv:2507.05515v4 Announce Type: replace Abstract: Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. In this work, we explore LEGO Co-builder, a hybrid benchmark combining real-world LEGO assembly logic with programmatically generated multimodal scenes. The dataset captures stepwise visual states and procedural instructions, allowing controlled evaluation of instruction-following, object detection, and state detection. We introduce a unified framework and assess leading VLMs such as GPT-4o, Gemini, and Qwen-VL, under zero-shot and fine-tuned settings. We also evaluated the framework using a reasoning-focused model, GLM-4.1
קרא במקור המקורי