יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כל דינים אינם שווים: חשוב מחדש על יציבות פסקי דין של LLM

All Verdicts are Not Equal: Rethinking LLM Judge Reliability
אנו מציגים סקירה עמוקה של יציבות פסקי דין של LLM, כולל ניתוח של שסתומים ופגמים שונים.
תקציר מקורי באנגליתarXiv:2610.12083v1 Announce Type: cross Abstract: LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formali
קרא במקור המקורי