כתבה
arXiv cs.CL ·
בדיקת אודיו גנרטיבי
Auditing generative audio calls for known-task audio-llm evaluation
חוקרים בדיקת אודיו גנרטיבי להערכת LLM. המחקר משווה נכונות עם ובלי קריאות אודיו גנרטיביות. התוצאות מראות כי הנכונות לא משתפרת משמעותית עם קריאות אודיו גנרטיביות.
תקציר מקורי באנגליתarXiv:2608.27817v4 Announce Type: replace-cross Abstract: Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call se
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית