כתבה
arXiv cs.AI ·
LAM יעיל עם העדפות אנוש
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
חוקרים בדקו 10 שיטות לבחירת תת-קבוצות קטנות להערכת מודלים קוליים גדולים. הם מצאו שתת-קבוצות קטנות יכולות לחזות העדפות אנושיות טוב יותר מאשר הבחינה המלאה.
תקציר מקורי באנגליתarXiv:2605.00022v2 Announce Type: replace-cross Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy. Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benchmark scores. To understand how well these scores align with what practitioners ultimately care about, user satisfaction, we collect 776 human preference ratings from realistic voice assistant conversations, finding that both subsets and full
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית