יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרופילי דיוק מותנים: אבחון LLM Judges

Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
פותחים פרופילי דיוק מותנים (CAP) לאבחון LLM Judges. המחקר מראה כי CAP יכול לשמש לאבחון LLM Judges בתנאים שונים. הפרופילים מספקים בסיס פעולה מעשי יותר מאשר דיוק מצטבר.
תקציר מקורי באנגליתarXiv:2610.09229v1 Announce Type: new Abstract: LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging b
קרא במקור המקורי