יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מחזון לשפה: חקר זרימת מידע גורמית בקבלת החלטות רב-מודאלית

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
חוקרים את זרימת המידע הגורמית בין חזון לשפה בקבלת החלטות רב-מודאלית. המחקר בוחן את הדרך בה מידע חזותי משפיע על החלטות שפה. התוצאות מראות כי מידע חזותי משולב בעיקר כאשר המודל עובד על אפשרויות תשובה.
תקציר מקורי באנגליתarXiv:2609.05149v1 Announce Type: new Abstract: Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semanti
קרא במקור המקורי