כתבה
arXiv cs.LG ·
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
תקציר מקורי באנגליתarXiv:2603.00039v2 Announce Type: replace Abstract: LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders -- such as verbosity, stylistic preferences, or training artifacts -- causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weigh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית