כתבה
arXiv cs.CL ·
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
תקציר מקורי באנגליתarXiv:2506.14766v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) frequently hallucinate by over-committing to spurious visual cues. Prior remedies-Visual and Instruction Contrastive Decoding (VCD, ICD)-mitigate this issue, yet the mechanism remains opaque. We first empirically show that their improvements systematically coincide with redistributions of cross-modal attention. Building on this insight, we propose Attention-Steerable Contrastive Decoding (ASCD), which directly steers the attention scores during decoding. ASCD combines (i) positive steering, which amplifies automatically mined text-centric heads-stable within a model and robust across domains-with (ii) negative steering, which dampens on-the-fly identified critical visual tokens. The method in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית