כתבה
arXiv cs.AI ·
NormViz: מבחן ופלטפורמה להבנת ראייה מולטימודלית בתרבויות גלובליות
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
מבחן ופלטפורמה חדשים להבנת ראייה מולטימודלית בתרבויות גלובליות. המבחן, NormViz-Bench, כולל 3,268 זוגות תמונות חד-צדדיות, והפלטפורמה, NormViz-Train, כוללת 64,000 תמונות עם הסברים. המבחן נועד לבחון את יכולת המודלים להבין ולהסביר תרבויות ומנהגים שונים.
תקציר מקורי באנגליתarXiv:2609.06831v1 Announce Type: new Abstract: AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irreleva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית