כתבה
arXiv cs.CL ·
ראות מה צריך לשמוע: ניתוח ותיקון שבילי קרוס-מודלי ב-LLMs אומני-מודלי
Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
ניתוח ותיקון שבילי קרוס-מודלי ב-LLMs אומני-מודלי. נמצא כי LLMs תלויים יתר על המידה במידע חזותי כשמתייחסים לשאלות אודות קול. פותחה טכניקה לתיקון זה.
תקציר מקורי באנגליתarXiv:2609.36798v1 Announce Type: cross Abstract: Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית