כתבה
arXiv cs.AI ·
Detect Before You Leap: Mirage Detection in Vision-Language Models
תקציר מקורי באנגליתarXiv:2606.00435v5 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can produce confident answers without relevant visual evidence, a failure mode known as mirage (Asadi et al., 2026). We study pre-release mirage detection: deciding whether a VLM answer should be released or withheld. Our model-agnostic method, Text-Conditioned Layer-wise Internal Alignment (TC-LIA), tracks question-image alignment across the layers of a frozen CLIP ViT-H/14 encoder, summarizing patch-text alignment by final similarity, late-layer top-k alignment, early-to-late gain, and slope. TC-LIA is training-free at deployment with fixed projections and scoring weights, without any label-specific training, and delivers strong detection independently. Additionally, when combined with blank/noise det
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית