יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לא ידיעת פעולה, אלא קישור לחלק: מקום הבלימה בתחזית יכולת-תמונה

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction
מודלי יכולת-תמונה נתקלים בקשיים במשימות פעולה נמוכ-רמה. התוצאות מצביעות על קשיים בקישור לחלקים של פריטים, ולא בחוסר ידיעת פעולה. זה נצפה בכל שלושת משפחות המודלים. התוצאות נשמרו גם לאחר שהפריט נקרא בשםו. זה קשור לקשיים באתגר של זיהוי חלקים.
תקציר מקורי באנגליתarXiv:2609.13225v1 Announce Type: cross Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulated objects we asked eight models, spanning three developers, what motion a robot should apply. Under an open prompt, push was produced once in 64 evaluations where it was correct, despite being correct for 8 of 19 objects and appearing in the offered label set every time. Inspecting the outputs showed why: models described a different part than the one being scored, e.g. explaining how to pick up a camera rather than press its
קרא במקור המקורי