כתבה
arXiv cs.CL ·
The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
תקציר מקורי באנגליתarXiv:2609.35860v1 Announce Type: new Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|\rho|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית