כתבה
arXiv cs.LG ·
Unexplored flaws in multiple-choice VQA make benchmarking unreliable
תקציר מקורי באנגליתarXiv:2511.22341v2 Announce Type: replace-cross Abstract: Previous works identify sensitivity to option order as a key issue in multiple-choice VQA (MC-VQA) evaluation and propose protocols to mitigate this effect. We show that such mitigation is insufficient to ensure the validity of MC-VQA as a reliable benchmark for Multimodal Large Language Model (MLLMs): performance remains highly sensitive to semantically neutral prompt format choices that are not controlled by current benchmarks. In a large-scale study spanning seven MLLMs and five MC-VQAs datasets, we find frequent rank reversals even under order-invariant evaluation. These reversals arise when we systematically vary option ID sets, delimiters, and separators, yielding 48 semantically equivalent prompt formats. Mechanistic analyses
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית