יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אשליות במודלים גדולים

When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
מודלים גדולים ממשיכים להציג אשליות, כלומר תגובות נראות אך שגויות. מחקר זה מראה כי קורלציות מזויפות בנתוני האימון גורמות לאשליות אלו, שעמידות לשיטות גילוי קיימות. המחקר בדק מודלים כגון GPT-5.
תקציר מקורי באנגליתarXiv:2511.07318v3 Announce Type: replace Abstract: Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and propri
קרא במקור המקורי