כתבה
arXiv cs.AI ·
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
תקציר מקורי באנגליתarXiv:2607.04383v4 Announce Type: replace-cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of these paradigms lies the task of Open-Vocabulary Audio Event Grounding: predicting all time intervals of a target sound event described by an arbitrary natural language query. Progress is bottlenecked by data scarcity: no large-scale resource provides open-vocabulary onset/offset supervision, and manual temporal annotation is prohibitively expensive. To address this, we introduce Auto-AEG, a scalable pipeline that constructs such supervision by automatic data construction and model fine-tuning. It
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית