כתבה
arXiv cs.CL ·
שיפור איכות שמע במודלים של שפה-קול
Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
LTM-AE הוא אלגוריתם חדש לשיפור איכות שמע במודלים של שפה-קול. הוא משתמש בזיכרון ארוך-טווח כדי לשפר את היכולת לזהות קולות מסוימים ברקע רעש. האלגוריתם נבדק על מודל Qwen2-Audio והראה שיפור משמעותי בדיוק זיהוי הקולות.
תקציר מקורי באנגליתarXiv:2609.36577v1 Announce Type: cross Abstract: Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית