כתבה
arXiv cs.CL ·
אני לא מתגעגע אליך, אבל אני כן: אמינות ההסבר העצמי של חסרונם של מודלים ויז'ואלי-לשוניים
I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models
מודלים ויז'ואלי-לשוניים נוטים להעריך יתר על המידה את המספיקות של הראיות הזמינות. המחקר בדק עשרה מודלים ומצא כי הם מעריכים בטעות את השפעת החסרונות והראיות הזמינות.
תקציר מקורי באנגליתarXiv:2609.07596v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support, whether restoring a missing modality would change their answer, and whether the available evidence is sufficient; we then execute the corresponding intervention and compare these claims with realized behavior. We evaluate ten VLMs spanning open-weight and proprietary models across four tasks covering mixed, redundant, and unique modality regimes. We find a sy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית