כתבה
arXiv cs.LG ·
Accounting for Bias Enables Sustainable LLM Evaluation
תקציר מקורי באנגליתarXiv:2609.31184v1 Announce Type: cross Abstract: LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית