יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פירוק אי-ודאות LLM-Judge לסימון תוויות מומחים

Decomposing LLM-Judge Uncertainty to Target Expert Labels
חוקרים פיתחו מודל בייסיאני קטן לפירוק אי-ודאות LLM-Judge. המודל מאפשר לזהות היכן השופט LLM אינו בטוח ולהפנות תוויות מומחים לשם. המחקר מראה כי הגישה החדשה מסלקת 83% יותר שגיאות מאשר אי-ודאות כוללת.
תקציר מקורי באנגליתarXiv:2609.06444v2 Announce Type: cross Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for the
קרא במקור המקורי