כתבה
arXiv cs.LG ·
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
תקציר מקורי באנגליתarXiv:2610.07755v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית