כתבה
arXiv cs.LG ·
GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them
מודלי VLMs נשפעים מסימני סמכות בתמונות, גם אם ניתנה ההוראה הספציפית להתעלם מהם.
תקציר מקורי באנגליתarXiv:2609.06058v1 Announce Type: cross Abstract: As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct conflict: a student answer, which should attract sycophantic agreement, and a conflicting answer attributed to a peer, teacher, or official answer key, which should attract authority-based deference. Models produce free-form answers while being explicitly instructed to solve independently and ignore all student answers, feedback, and grading marks. We test the models on 60 synthetic real-world trade-off scenarios. Five neutral trials establish a stabl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית