כתבה
arXiv cs.AI ·
Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception
תקציר מקורי באנגליתarXiv:2609.15215v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative formulation lacks explicit correspondence between the timestamp predictions and fine-grained acoustic evidence, limiting the precision and reliability of temporal localization. To address this issue, we augment the LALM with a dedicated frame-level grounding model while leveraging its semantic modeling capability to represent the event query. Specifically, the frozen LALM encodes the event query with audio as context, and the g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית