יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MMAC: בנק אודיו גדול לתיאורי קפצ'ן

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
MMAC הוא בנק אודיו גדול שמכיל 5,638 קליפים, המכסה 6 קטגוריות ו-15 ממדים. MMAC בודק את יכולת המודלים לתאר קפצ'ן אודיו.
תקציר מקורי באנגליתarXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the refere
קרא במקור המקורי