כתבה
arXiv cs.CL ·
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
תקציר מקורי באנגליתarXiv:2509.25160v3 Announce Type: replace-cross Abstract: Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs. It remains unclear whether VLMs can reason mathematically when information must be perceived and inferred from images rather than read from explicit symbols. We introduce GSM8K-V, a benchmark transforming GSM8K into multi-image sequences with semantic equivalence preserved. By mapping text-based problems into visual form via an automated pipeline and human verification, we curate 1,319 high-quality samples. In GSM8K-V, quantities must be extracted through visual perception, and reasoning chains must be reconstructed by integrating implicit cues across scenes. Evaluation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית