כתבה
arXiv cs.CL ·
SWORD: בדיקת שגיאות עובדתיות במודלים רב-לשוניים
SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection
SWORD היא בדיקה להערכת יכולתם של מודלים רב-לשוניים לזהות שגיאות עובדתיות. היא משתמשת בנתוני Wikidata כדי ליצור הצהרות שגויות בשמונה שפות. התוצאות מראות שמודלים מסוימים מתקשים יותר עם שפות מסוימות, כגון שפות אסיאתיות.
תקציר מקורי באנגליתarXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, ranging from random entity substitutions to semantically plausible property-based selections. Our distortion-based evaluation surfaces two critical insights that remain entirely obscured by conventional benchmarks. First, models counter
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית