יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בדיקת VLM-גיידד אימג' סלקשן

When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts
מחקר זה בודק את היכולת של מודלים ויז'ואל-לשוניים (VLMs) לבחור את התמונה הטובה ביותר מבין מספר תמונות שנוצרו. המחקר מראה כי VLMs עם 4B פרמטרים מסתמכים יותר מדי על התמונה הראשונה שמוצגת, וכי החלפת הסדר של התמונות יכולה לשנות את הבחירה. VLMs עם 8B פרמטרים מראים ביצועים טובים יותר, אך עדיין יש להם נטייה לבחור תמונות מסוימות.
תקציר מקורי באנגליתarXiv:2610.01243v1 Announce Type: cross Abstract: Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is inf
קרא במקור המקורי