יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

FPCO-Dialog: בדיקת שגיאות ושיתוף פעולה

FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
FPCO-Dialog הוא בדיקת ביצועים למודלים של שפה-תמונה. הוא בודק את היכולת לתקן שגיאות ולשתף פעולה בדיאלוגים. הבדיקה כוללת 1,080 תמונות ו-10,800 שאילתות.
תקציר מקורי באנגליתarXiv:2609.03331v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dialogue turns. We introduce FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises. FPCO-Dialog contains 1,080 images and 10,800 question turns, stratified by visual complexity, object category, and false-premise class, and uses a 10-turn protocol in which a correct dialogue prefix is followed by repeated false-premise referring expressions. We evaluate 20 commercial and open-source VLMs with a model-agnostic protocol an
קרא במקור המקורי