כתבה
arXiv cs.AI ·
ראיות חדשות, אותה בחירה: בדיקת בחירת ניסויים פיזיים במודלים של שפה וראייה
New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models
חוקרים בדקו את היכולת של מודלים של שפה וראייה לבחור ניסויים פיזיים. המודלים קיבלו תמונה מניסוי פיזי ונדרשו לענות על שאלה או לבחור ניסוי נוסף. התוצאות הראו כי המודלים אינם מסוגלים לקבל החלטות מדויקות.
תקציר מקורי באנגליתarXiv:2609.11022v1 Announce Type: cross Abstract: A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית