יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שלושה דרכים שתאוריית המבחן הקלאסי יכולה להטעות בקריאת דיונים של LLM

Three Ways Classical Test Theory Can Mislead About LLM Judges
תאוריית המבחן הקלאסי יכולה להטעות בקריאת דיונים של LLM בשלושה דרכים. המחברים חקרו את השפעתה של תאוריית המבחן הקלאסי על קריאת דיונים של LLM ומצאו שהיא יכולה להטעות בשלושה דרכים שונים. הם גם חקרו את השפעתה של תאוריית המבחן הקלאסי על קריאת דיונים של LLM ומצאו שהיא יכולה להטעות בשלושה דרכים שונים.
תקציר מקורי באנגליתarXiv:2609.29709v2 Announce Type: replace-cross Abstract: Evaluations that use a large language model (LLM) as a judge have begun to borrow reliability statistics from classical test theory and its extensions. We examine three such statistics that need one administration and no gold labels. None of them can isolate the judge, because one judge under one prompt supplies no variance component of its own. Claude Haiku 4.5 judged 210 constructed short answers against ten-element checklists. On the 180 with parsed verdicts, the Kuder-Richardson coefficient (KR-20) came out at 0.5223 on the judge's verdicts and 0.5231 on error-free gold verdicts. In simulation, bank design alone moves KR-20 from 0.01 to 0.68 at the judge's measured 4.72% error rate. The dependability index $\Phi(\lambda)$, a rat
קרא במקור המקורי