יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעקב אחר היקשרות אודיו ובחירת תשובות במודלים LLM אודיו

Tracing Audio Grounding and Answer Selection in Audio LLMs
מודלים LLM אודיו משפרים את הבנת האודיו, אך עדיין יכולים לנבא תשובות מרמזים טקסטואליים. מחקר זה בודק מה קורה בתוך המודל כדי שהאודיו יקבע את התשובה. התוצאות מראות שאימון המודל על נתונים שאינם יכולים להיות מוסקים מטקסט בלבד משפר את הביצועים.
תקציר מקורי באנגליתarXiv:2609.04637v1 Announce Type: cross Abstract: Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve performance, but what changes within the model remains unclear. In this paper, we ask what must happen inside the model for the audio to actually determine the answer. Our findings are threefold. (1) Replacing the audio with silence or unrelated audio causes substantially larger performance degradation in the trained model than in the pretrained model. (2) Acoustic information most strongly shapes the model's representations o
קרא במקור המקורי