יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בחירות דומות, תשומת לב שונה: קשרים מודליים-מודליים בבני אדם ובמודלי תצוגה-שפה

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models
חוקרים חקרו קשרים מודליים-מודליים בבני אדם ובמודלי תצוגה-שפה. הם גילו שמודלים גדולים תואמים בחירות, אך לא תשומת לבם. תרגול מודלים קטנים על בחירות אדם הביא לתואמות בחירות, אך לא לתשומת לב.
תקציר מקורי באנגליתarXiv:2609.36475v1 Announce Type: cross Abstract: Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fi
קרא במקור המקורי