כתבה
arXiv cs.CL ·
Decomposing LLM-Judge Uncertainty to Target Expert Labels
תקציר מקורי באנגליתarXiv:2609.06444v3 Announce Type: replace Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית