כתבה
arXiv cs.AI ·
UniAE-MoE: מקודד אודיו אחוד
UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
UniAE-MoE הוא מקודד אודיו אחוד שנועד לדגמי אודיו בתחומים שונים. הוא משלב טכנולוגיות מ-Qwen2-Audio ו-Audio-Flamingo 3, ומשיג ביצועים טובים במגוון משימות אודיו.
תקציר מקורי באנגליתarXiv:2609.39199v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית