כתבה
arXiv cs.CL ·
האם מודלים רב-מודאליים רואים לפני שהם קוראים?
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
חוקרים גילו כי מודלים רב-מודאליים כושלים כאשר טקסט מבטל ראיות חזותיות. המחקר בדק את התופעה הזו במודלים שונים, כולל GPT-5.1 ו-Gemini. התוצאות הראו כי המודלים נוטים להעדיף טקסט על פני ראיות חזותיות.
תקציר מקורי באנגליתarXiv:2609.00067v2 Announce Type: replace Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית