יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

JEV-שופט: מקבל כאשר בטוח, מעלה כאשר לא בטוח

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
JEV-שופט הוא שיטה להערכת LLM, המשתמשת בשופט החלטה בלבד, שמחזיר הסתברות לייחוס תווית, ולא טקסט. השופט מחליט האם לקבל את החלטתו או להעלות לשופט הסבר. השיטה הראתה תוצאים טובים בהשוואה ל-GPT-6.
תקציר מקורי באנגליתarXiv:2609.26550v4 Announce Type: replace Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than
קרא במקור המקורי