כתבה
arXiv cs.LG ·
כאשר הטיה תואמת: כיצד קורלציות שקריות פוגעות בדטקטור של הללומים
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
מודלי הללומים (LLMs) חווה הללומים עקב קורלציות שקריות בנתוני הלמידה. המחקר חושף כי קורלציות אלה גורמות להללומים שאין להם כל קשר למקור, ואינם נפגעים על ידי שיטות דטקטור שקיימות כיום.
תקציר מקורי באנגליתarXiv:2511.07318v3 Announce Type: replace-cross Abstract: Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית