כתבה
arXiv cs.CL ·
בחינת עמידות LLM כשופט: הבחנה פסיכומטרית של קשיי שיפוט
Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty
במאמר זה, חוקרים בחנו את עמידות LLMs כשופטים, וגילו שהם עשויים להפריט קשיי שיפוט שונים מאלו של בני אדם. המחקר עוסק בבחינת LLMs כשופטים, ובפרט בבחינת עמידותם. החוקרים גילו ש-LLMs עשויים להפריט קשיי שיפוט שונים מאלו של בני אדם, ושהם עשויים להיות יעילים יותר בקשיי שיפוט מסוימים. המחקר יכול לשפר את יעילות LLMs כשופטים, ולהפחית את כשלי השיפוט שלהם.
תקציר מקורי באנגליתarXiv:2610.02877v1 Announce Type: new Abstract: Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית