כתבה
arXiv cs.LG ·
מה ששמע ה-MLLM? גירוי ספקטרלי-זמני להסברה של MLLM לאודיו
What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability
פרסום: גירוי ספקטרלי-זמני להסברה של MLLM לאודיו. STAG, פלטפורמה חדשה, מספקת הסברה טובה יותר לאודיו-MLLM.
תקציר מקורי באנגליתarXiv:2609.12663v1 Announce Type: cross Abstract: Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support for each generated token using target-token-specific vocabulary projections of the encoded audio representations, measures frequency-band relevance t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית