יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

היקשרות מציאותית של תצפיתן

Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models
חוקרים בדקו יכולתם של מודלים רב-תחומיים להבין יחסים מרחביים מנקודת מבטו של דובר. הם פיתחו את POVBench, מאגר נתונים לבדיקת היכולת הזו. התוצאות מראות כי אפילו מודלים מתקדמים מתקשים לאתר מטרות לא נראות או לא מוגדרות היטב.
תקציר מקורי באנגליתarXiv:2609.06880v1 Announce Type: cross Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated perspective. Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. Recent vision-language models (VLMs) appear capable of spatial reasoning, but their ability to infer a speaker's viewpoint from contextual cues and interpret situated spatial relations from that viewpoint remains unclear. We call this capability contextual observer grounding. To study this capability, we construct the Point-of-View Benchmark (POVBench), a dataset of 3D scenes and queries that disentangles Inferred, Stated, and Given forms of observer grounding in natural embodied comm
קרא במקור המקורי