כתבה
arXiv cs.AI ·
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
תקציר מקורי באנגליתarXiv:2607.15176v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית