כתבה
arXiv cs.CL ·
Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
תקציר מקורי באנגליתarXiv:2609.37577v1 Announce Type: new Abstract: Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley--Terry geometry underlying pairwise aggregation, each proxy is dominated by close-rank-gap pairs, where inconsistency is information-theoretically expected and individual verdicts contribute little to the aggregate ranking; far-gap pairs
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית