יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

לחץ דיסקורסיבי: עבודה עם מבנה המידע במודלי תקשורת-תמונה

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
מודלי תקשורת-תמונה נתקלים בקושי לבטא תוכן תמונתי בצורה המתאימה לדיסקורס. המחקר חושף כי המודלים נוטים להתקרר ולהפגין רגישות יתר למבנה המידע.
תקציר מקורי באנגליתarXiv:2605.28346v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they express such content in a discourse-appropriate form. We address this research gap using information structure (IS), testing whether VLMs distinguish discourse-old Topics from discourse-new Foci in visually grounded question answering. We exploit Hungarian, a language in which Topic and Focus map onto dedicated syntactic positions, making IS choices observable in text. Comparing six VLMs with human participants, we find that models produce IS-relevant constructions, but over-regularise this sensitivity. Under the interacting pressures of discourse status, grammatical role (preference for subje
קרא במקור המקורי