כתבה
arXiv cs.AI ·
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
תקציר מקורי באנגליתarXiv:2601.08654v3 Announce Type: replace-cross Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the same criteria inconsistently, produce score attributions that are difficult to audit, and map judgments poorly onto human scoring scales. We define this challenge as criteria transfer: translating human rubric intent into a stable, auditable inference-time scoring protocol. We introduce Rulers, which locks a task-level rubric specification, executes it through structured, evidence-grounded judgments, and calibrates the resulting signals to human score boundaries. Across four rubric-governed benchmarks and multiple frozen backbone models, Rulers achieves stronger agreement with human score
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית