כתבה
arXiv cs.LG ·
Evaluating Rubric Generation with Interventional Transfer
תקציר מקורי באנגליתarXiv:2610.10809v1 Announce Type: new Abstract: Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית