כתבה
arXiv cs.LG ·
AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models
תקציר מקורי באנגליתarXiv:2610.06949v1 Announce Type: cross Abstract: Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3\% of the base model's parameters and plugs into any audio encoder--
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית