כתבה
arXiv cs.AI ·
JudgeProfile: הבנה ושליטה בסובייקטיביות של חברי מושבעים ב-LLM
JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
חוקרים סובייקטיביות של חברי מושבעים ב-LLM ומציגים תוכנית להבנה ושליטה בה.
תקציר מקורי באנגליתarXiv:2609.36705v2 Announce Type: replace Abstract: LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize eac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית