כתבה
arXiv cs.CL ·
הגבלות של מדדי תקציב חינם לכתיבה: כשלים כמדדים וכתגמילים לאופטימיזציה
The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
מדדים חינם לכתיבה נכשלים לעקוב אחר העדפות האדם בביקורת טקסט-לדיבר. המאמר עורכב על תצפיות של כ-6 קורפוסים של קליפים טקסט-לדיבר, ובהם קליפים עם תקלות וקליפים חדשים. המדדים נבדקו על ידי תצפיות של קליפים שנבחרו על ידי האדם. התוצאות הראו שהמדדים נכשלו לעקוב אחר העדפות האדם. נראה שהפתרון היחיד לבעיה הוא להשתמש במדדים חינם שונים, ולאחר מכן להשתמש במדדים של האדם.
תקציר מקורי באנגליתarXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingly adopted as reward signals for preference optimization. Both roles presuppose that the predicted score tracks human preference. In this work, we test this assumption across six human-rated corpora spanning the quality range from artifact-rich to defect-free TTS, evaluating each predictor on a pairwise task that asks whether the clip it scores higher is the clip listeners prefer, and we subject interpretable prosodic and signal-processing features to the same protocol. When one clip carries audible defects the predictors tend to agree with listeners. Once both clips are clean, no single predict
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית