כתבה
arXiv cs.AI ·
מודלים רב-מודאליים גדולים מדווחים על תמונות דו-יציבות
(How) Do MLLMs Report Bistable Images Like Humans?
חוקרים בדקו את התנהגותם של מודלים רב-מודאליים גדולים (MLLMs) בהקשר לתמונות דו-יציבות, כגון תמונת הברווז-ארנב. התוצאות הראו שהמודלים מדווחים על התמונות באופן דומה לבני אדם, עם השפעה של רמזים חזותיים ולשוניים.
תקציר מקורי באנגליתarXiv:2609.13254v1 Announce Type: cross Abstract: Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consiste
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית