כתבה
arXiv cs.LG ·
Understanding the Effects of Distractors on Reasoning Vision-Language Models
תקציר מקורי באנגליתarXiv:2511.21397v3 Announce Type: replace-cross Abstract: How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distractors can intensify inverse scaling, causing models to reason longer but less effective reasoning traces. In this work, we investigate whether similar phenomena arise in multimodal settings. We introduce Idis (Images with distractors), a visual question-answering dataset that systematically varies distractors along semantic and numerical dimensions. Our analyses reveal that visual distractors affect reasoning VLMs in a fundamentally different way from textual distractors: although inverse scaling still emerges, visual distractors reduce accuracy without in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית