כתבה
arXiv cs.CL ·
בחינת VQA בדגמי שפה-תצלום באמצעות עקרונות שיתופי
Evaluating VQA in Vision Language Models using Cooperative Principles
במאמר זה, נבחן את יכולתם של דגמי שפה-תצלום לפתור שאלות VQA כאשר השאלות חורגות מעקרונות שיתופי. נמצא כי דגמי ChatGPT, Claude, Gemini ו-Llama מציגים יכולת פחות טובה במקרים אלה. כמו כן, נמצא כי דגמי שפה-תצלום עצמם פועלים באופן שונה מאדם בפתרון סטיות של עקרונות שיתופי.
תקציר מקורי באנגליתarXiv:2610.02878v1 Announce Type: new Abstract: We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית