כתבה
arXiv cs.AI ·
CLM-אירוע: הערכת מודל קונטרסטיבי פתוח
CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks
מודל קונטרסטיבי פתוח הוערך על בנקי מבחן ציבוריים. המודל השיג תוצאות נמוכות, דומות להטלת מטבע. מודלים אחרים עם אותו מספר פרמטרים השיגו תוצאות גבוהות יותר.
תקציר מקורי באנגליתarXiv:2610.07177v1 Announce Type: cross Abstract: An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית