יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מעניקים עדיפות לבני אדם: כלי חדש לבדיקת דגמי קול

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
במאמר זה, המחברים פיתחו כלי חדש לבדיקת דגמי קול, המתמקד בעדיפות האדם. הם ניסו 10 שיטות שונות לבחירת תת-קבוצות ומצאו ש-50 דגימות (0.3% מהנתונים) יכולות להגיע לתוצאות של 0.93 קורלציה עם תוצאות הבנצ'מרק המלא. הם גם קבעו שהכלי שלהם, HUMANS, מגיע ל-0.98 קורלציה עם תוצאות הבנצ'מרק המלא, והוא יכול לחסוך כ-99% מהנתונים.
תקציר מקורי באנגליתarXiv:2605.00022v2 Announce Type: replace Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy. Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benchmark scores. To understand how well these scores align with what practitioners ultimately care about, user satisfaction, we collect 776 human preference ratings from realistic voice assistant conversations, finding that both subsets and full bench
קרא במקור המקורי