יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

האם ניתן להסמיך פסנתרים LLM: מחקר של התלות ביכולות ובייסים ניתנים לתיקון והצבעה קבוצתית

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
במחקר זה נחקרו הבייסים של פסנתרים LLM והוצעה שיטה להצבעה קבוצתית שמציעה הסמכה ניתנת לתיקון. המחקר כלל שישה מודלים שונים ובחן את יכולתם לזהות טעויות. התוצאות הראו שהשיטה החדשה מציעה הסמכה ניתנת לתיקון ומשפרת את דיוק ההצבעה.
תקציר מקורי באנגליתarXiv:2609.12002v1 Announce Type: new Abstract: LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use cases. Across four benchmarks and six models (36 judge-examinee pairs), we show that a model's task accuracy strongly predicts its judging accuracy (Pearson $r \geq 0.90$ on most models) and inversely predicts its directional bias ($r \leq -0.83$), but that accuracy alone does not ensure fair evaluation: more capable examinee models consistently receive more lenient judgments from all judges ($r \geq 0.83$). To address this, we prop
קרא במקור המקורי