כתבה
arXiv cs.LG ·
יכולת אחת או רבות? בדיקות מבניות וחיזויות של תקפות בנצ'מרק
One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI
מחקר זה בודק את תקפותן של בנצ'מרקים כלכליים להערכת יכולות AI. הממצאים מראים כי הבנצ'מרקים הכלכליים אינם מודדים יכולת נפרדת, אלא חלק מיכולת כללית. המחקר מציע פרוטוקול לבנאי בנצ'מרקים.
תקציר מקורי באנגליתarXiv:2608.29420v2 Announce Type: replace Abstract: Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which 103 configurations carry all three sparsely scored economic benchmarks and 96 carry all twelve; four hypotheses and their thresholds were fixed before analysis, and every deviation from the plan is reported. The first factor of a three-factor extra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית