כתבה
arXiv cs.CL ·
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
תקציר מקורי באנגליתarXiv:2607.14480v2 Announce Type: replace Abstract: LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource language
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית