כתבה
arXiv cs.AI ·
VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics
תקציר מקורי באנגליתarXiv:2609.37287v1 Announce Type: cross Abstract: Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית