כתבה
arXiv cs.CL ·
עד כמה רגישים דירוגי LLM לשינויים בבסיס הנתונים?
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
דירוגי LLM עשויים להיות תלויים בהרכב הבסיס הנתונים. נמצא כי דירוגי LLM של חמשה בסיסי נתונים שונים היו רגישים לשינויים בבסיס הנתונים. זה חשוב להתחשב בכך כאשר מדירגים LLM.
תקציר מקורי באנגליתarXiv:2609.00482v2 Announce Type: replace Abstract: Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF). These items are used to score models in the other half with frozen weights that preserve the benchmark's composition across metadata-defined item groups and easiness strata. Equally short matched-random subtests provide a baseline for variation due to item subsampling. Full-benchmar
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית