כתבה
arXiv cs.AI ·
הפער הניכור: הטבעות הסודיות בהכרזת ההלצות של מודלי שפה
The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
מחקר זה חוקר את הטבעות הסודיות בהכרזת ההלצות של מודלי שפה, ומצא כי קיימות הבדלים ניכורים בין המודלים. המחקר חושף פערים באמינות הדטקציה של ההלצות, ומציע דרכים לשפר את האמינות של המודלים.
תקציר מקורי באנגליתarXiv:2609.35860v1 Announce Type: cross Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|\rho|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dis
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית