כתבה
arXiv cs.CL ·
JEV נגד LLMs כשופטים: זולים, מהירים, ולא צודקים באותם הנקודות
JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places
LLMs נגד JEV כשופטים: זולים, מהירים, ולא צודקים באותם הנקודות. המחקר מציע כי קטעי עבודה של LLMs יכולים להיות חלופה ל-JEV, אך שני הסוגים של שופטים טועים באותם הנקודות.
תקציר מקורי באנגליתarXiv:2609.29769v2 Announce Type: replace Abstract: LLM judges score outputs against rubrics well enough to have become the norm, both in benchmarks and as rewards for training. Jev, a classifier-like alternative its creators call a "decision model", returns probabilities over permitted answers with a calibrated confidence score, which LLM judges do not natively provide. We compare Jev with three flash-tier LLM judges on nine panels from seven benchmarks with human judgments, giving every judge identical criterion texts. The LLM judges run in two setups: holistically, reading a whole rubric at once as Jev does, and one criterion at a time. Jev can often stand in for them. They cost 16 to 325 times as much and take 28 to 350 times as long, yet in each setup Jev's accuracy differs significan
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית