כתבה
arXiv cs.AI ·
מעבר מאיכות הפסיקה ליועצות רב-ערך: חשיבה מחדש על LLM-כפסיק
From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks
במאמר זה, חוקרים חושבים מחדש על שימוש ב-LLM כפסיק לצורך הערכת תגובות למשימות פתוחות. הם חוקרים את איכות הפסיקה ואת היועצות המוצאות, ומצאו כי לא תמיד יש קשר בין האיכות ליועצות.
תקציר מקורי באנגליתarXiv:2609.37145v1 Announce Type: new Abstract: LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided rev
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית