כתבה
arXiv cs.AI ·
PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
תקציר מקורי באנגליתarXiv:2609.36086v1 Announce Type: cross Abstract: Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADM\'E, a data synthesis method that generates reliable criterion-ba
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית