כתבה
arXiv cs.LG ·
מעקב אחר הקשר בין הקלטה ובחירת התשובה ב-LLMs לשמע
Tracing Audio Grounding and Answer Selection in Audio LLMs
LLMs לשמע יכולים לגרום להצעות התשובה על פי קלט טקסט, אך הכשרה על נתוני שמע בלבד משפרת את התוצאות.
תקציר מקורי באנגליתarXiv:2609.04637v1 Announce Type: cross Abstract: Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve performance, but what changes within the model remains unclear. In this paper, we ask what must happen inside the model for the audio to actually determine the answer. Our findings are threefold. (1) Replacing the audio with silence or unrelated audio causes substantially larger performance degradation in the trained model than in the pretrained model. (2) Acoustic information most strongly shapes the model's representations o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית