יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעבר להחלטה: הערכה-מאוזנת של גבולות הזרקה-ביקורת

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails
במאמר זה נחקרה הערכה-מאוזנת של גבולות הזרקה-ביקורת במודלי תצוגה-שפה. המחברים פיתחו בנק אינסטרוקציות-צילום של 9,954 זוגות, כולל כותרות-התאמה, רשתות-זיכרון-מדויקות ופרטים-צלם-מתאימים. התוצאות הראו ש-2 מודלים עם דיוק-ממוצע זהה חושבים 9 פעמים פחות תצוגות-מאוזנות-הוכחה. המחברים פיתחו 2 פעולות-אינטרוונצ'ן שאינן-למידה, שמשפרות-התאמה-ללא-שינוי-פסק-דין.
תקציר מקורי באנגליתarXiv:2609.05535v1 Announce Type: cross Abstract: Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instruction-screenshot pairs with instruction-relative labels, pixel-exact evidence boxes, and matched image-side counterfactuals. Across six VLMs, two models with nearly identical average precision differ ninefold in Evidence-Aligned Detection (EAD), the fraction of attacks both detected and correctly localized. To test whether a verdict depends on the command cited as evidence, we replace the instruction with one that endorses th
קרא במקור המקורי