כתבה
arXiv cs.AI ·
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
תקציר מקורי באנגליתarXiv:2605.09883v3 Announce Type: replace-cross Abstract: As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores genuinely reflect robust visual understanding? We identify a pervasive vulnerability, the Cartesian Shortcut: models frequently discretize the orthogonal grid-based layouts prevalent in visual reasoning benchmarks into explicit textual coordinates, offloading reasoning from visual perception to text-based deduction. To re-evaluate visual reasoning when this shortcut is unavailable, we introduce Polaris-Bench, which re-formulates 53 visual reasoning tasks in Polar coordinate space with paired Cartesian counterparts that preserve task semantics, disrupting the orthogonal structure that mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית