כתבה
arXiv cs.AI ·
DiffWAM: A Fast and Efficient Navigation World Action Model
תקציר מקורי באנגליתarXiv:2609.39763v1 Announce Type: cross Abstract: Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית