כתבה
arXiv cs.AI ·
תשומת לב יעילה: תכנון מודלים לעקוב אחר תשומת לב אדם
Emergent Goal-Directed Attention in Large Vision-Language Models
מודלי תצפיות-שפה גדולים יכולים להעדיף מידע ראייתי כמו בני אדם, ללא הסברה של תשומת לב.
תקציר מקורי באנגליתarXiv:2609.05517v1 Announce Type: cross Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free viewing, leaving open whether goal-directed attention can emerge in systems without gaze supervision. We tested two off-the-shelf vision-language models (VLMs), Qwen3-VL-32B-Thinking and Gemma-4-26B-A4B-it, on 4,887 naturalistic scenes under visual-search and free-viewing instructions. Model predictions were compared with human fixations on the same images under corresponding tasks. Both models aligned more closely with human fixations under matching goals than under mismatched goals. This crossover persisted in target-absent scenes, where alignment could not be explained by simple visual groundi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית