כתבה
arXiv cs.AI ·
The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
תקציר מקורי באנגליתarXiv:2609.32964v2 Announce Type: replace Abstract: Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית