כתבה
arXiv cs.CL ·
From Routing Signals to Selective Review: Visual regrounding in MoE VLMs
תקציר מקורי באנגליתarXiv:2609.38111v1 Announce Type: new Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a ta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית