כתבה
arXiv cs.LG ·
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
תקציר מקורי באנגליתarXiv:2604.02543v2 Announce Type: replace-cross Abstract: As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We address this gap through a comprehensive empirical study of confidence calibration in VLMs, spanning three model families (Qwen3-VL, InternVL3, LLaVA-NeXT), three model scales (2B--38B), and multiple confidence estimation prompting strategies, across three medical visual question answering (VQA) benchmarks. Our study yields three key findings: First, overconfidence persists across model families and is not resolved by
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית