כתבה
arXiv cs.AI ·
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
תקציר מקורי באנגליתarXiv:2601.16520v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding, yet precise compositional spatial reasoning under geometric constraints remains underexplored. Existing benchmarks mainly assess coarse spatial relations and rarely support rigorous geometric verification or multiple valid solutions in constructive tasks. To address these limitations, we introduce TangramPuzzle, a benchmark comprising 668 validated configurations and 1,336 instances for evaluating compositional spatial reasoning under geometric constraints. We formulate the Tangram Construction Expression (TCE) to encode tangram configurations with machine-verifiable geometric specifications, together with a m
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית