כתבה
arXiv cs.AI ·
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
תקציר מקורי באנגליתarXiv:2605.18162v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have made striking progress, yet their spatial reasoning remains fragile. Models that answer an original input correctly can still fail under valid transformations with predictable answer mappings, revealing a gap between instance-level correctness and robust spatial reasoning. To address this, we propose Spatial Alignment via Geometric Evolution (SAGE), a self-evolving framework that improves robust spatial reasoning through geometric and linguistic duality operations. SAGE incorporates duality consistency into GRPO training, encouraging models to produce coherent answers across original and transformed inputs. SAGE co-evolves duality generation and solution, allowing the model to continually expose an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית