כתבה
arXiv cs.AI ·
אין זה מה שהתמונה מראה: תצלום רב-תכליתי נפגע מקונטקסט זר
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
תצלום רב-תכליתי נפגע מקונטקסט זר, וזה לא משפיע על התוצאות. חלק מהמודלים נפגעים יותר מאחרים.
תקציר מקורי באנגליתarXiv:2609.37863v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleadin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית