כתבה
arXiv cs.CL ·
CapGeo-Bench: Decoupling Visual Perception from Reasoning and Evaluating Geometric Understanding
תקציר מקורי באנגליתarXiv:2510.09302v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable success in difficult purely textual mathematical reasoning tasks, even advanced closed-source models such as GPT-o3 still struggle with geometric problems. This discrepancy motivates us to investigate the root cause: is the bottleneck of multimodal geometric reasoning rooted in reasoning itself, or in the perception of geometric information from diagrams? To answer this question, we conduct an exploratory experiment and find that providing high-quality captions consistently and substantially boosts performance of MLLMs, empirically validating the visual perception bottleneck in geometric reasoning. However, MLLMs' capabilities in visual geometric perception rema
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית