כתבה
arXiv cs.AI ·
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
תקציר מקורי באנגליתarXiv:2606.15753v4 Announce Type: replace Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-only or coordinate-augmented chain-of-thought, where entity references remain implicit and ambiguous. This may cause the reasoning process to decouple from visual evidence, entity references to drift across steps, and a causal disconnection between the reasoning trajectory and the final answer, with these problems further amplified in multi-view scenarios due to cross-view appearance changes. To address these issues, we propose Pinned Chain-of-Thought (PinCoT), a structured reasoning paradigm that pins every re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית