כתבה
arXiv cs.AI ·
MM-IFEval-Pro: בנק המבחנים המרכזי לביצוע פקודות במודלי תצוגה-שפה
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
בנק המבחנים MM-IFEval-Pro, המכיל 4 קטגוריות עיקריות ו-24 תת-קטגוריות, כולל תסכילות פקודות חד-פעמיות והסתה. הבנק כולל 52 תת-קטגוריות, כל אחת כוללת 3.0 תקנות ממוצע. MM-IFEval-Pro נועד לבחון את יכולת המודלים לבצע פקודות בשפה הסינית ובאנגלית.
תקציר מקורי באנגליתarXiv:2609.04859v1 Announce Type: new Abstract: As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual and safety-sensitive settings. To address these gaps, we present MM-IFEval-Pro, a multimodal instruction-following benchmark covering Chinese and English tasks as well as diverse instruction hijacking cases. MM-IFEval-Pro includes 4 major task categories and 24 subcategories and 8 instruction categories wi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית