כתבה
arXiv cs.CL ·
Look Before You Leap: Factual Decoding with Internal Attribution Signals
תקציר מקורי באנגליתarXiv:2609.15745v1 Announce Type: new Abstract: Hallucination remains a critical challenge in large language models (LLMs), where early factual errors compound through autoregressive generation in a snowballing effect that neither post-hoc correction nor weight-level intervention can effectively preempt. We propose DescaPE (DEcoding Signal Control Against Path Error-snowballing), a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at inference time. Through sliding-window MLP ablation, we identify a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps. We train a lightweight probe to approximate this signal from a single forward pa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית