יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עד כמה בחירות עיצוב LLM-שופט משנות?

How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models
חוקרים בדקו את השפעת בחירות עיצוב LLM-שופט על תוצאות. הם מצאו שבחירות אלו יכולות לשנות את התוצאות, אך ה-LLM-שופטים הם אמינים בכללי. המחקר השתמש במודלים כמו LLaMA.
תקציר מקורי באנגליתarXiv:2610.05094v1 Announce Type: cross Abstract: Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human groun
קרא במקור המקורי