כתבה
arXiv cs.LG ·
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
תקציר מקורי באנגליתarXiv:2609.00482v2 Announce Type: replace-cross Abstract: Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF). These items are used to score models in the other half with frozen weights that preserve the benchmark's composition across metadata-defined item groups and easiness strata. Equally short matched-random subtests provide a baseline for variation due to item subsampling. Full-be
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית