כתבה
arXiv cs.AI ·
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2607.25912v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $\pi_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of $\pi_0$. This enables the policy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית