כתבה
arXiv cs.AI ·
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
תקציר מקורי באנגליתarXiv:2608.18091v2 Announce Type: replace-cross Abstract: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narrative constraints selected from a shared pool, which carry no stylistic fingerprint yet retain a recoverable model-specific signature. Two experiments on this task yield complementary findings. Under
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית