כתבה
arXiv cs.AI ·
VIF-Bench: בדיקת תיאור נראה באמצעות הנחיות רב-מקוריות
VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
VIF-Bench היא בסיס פתוח לבדיקת תיאור נראה באמצעות הנחיות רב-מקוריות. הבסיס כולל 1,241 תפקידים שמטרתם לבדוק את יכולות המודלים בתיאור נראה. הבסיס כולל תפקידים שונים, כגון יצירת תמונות רב-מקוריות, תיאור נראה והשוואה בין תיאור נראה לתיאור טקסטואלי.
תקציר מקורי באנגליתarXiv:2609.37709v1 Announce Type: cross Abstract: Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g.,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית