יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בחינה מחודשת של הערכה מותאמת לאדם

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER
חוקרים בודקים מטריקות חדשות להערכת איכות תעתיק אוטומטי. הם מציגים את HATS-en, מאגר נתונים להערכה מותאמת לאדם. התוצאות מראות כי WER אינו המדד הטוב ביותר להערכת איכות תעתיק.
תקציר מקורי באנגליתarXiv:2609.21663v2 Announce Type: replace Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language model, layer, and pooling strategy. We find that WER agrees least with human judgment among all metrics tested, that the best-performing SemDist configurations achieve the highest overall agreement, ahead of CER and BERTScore, and that no single mode
קרא במקור המקורי