יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כאשר עקביות לא משמעות אמינות

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
חוקרים בדקו את העקביות והאמינות של שופטים מבוססי LLM, כולל LLaMA-3-8B ו-Qwen2.5-7B, בהשוואה לציונים של בני אדם. התוצאות הראו כי עקביות גבוהה לא בהכרח משמעות אמינות גבוהה.
תקציר מקורי באנגליתarXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each response is scored by nine human annotators and is evaluated three times by each LLM judge using the same rubric. We compare the judge scores with the averag
קרא במקור המקורי