כתבה
arXiv cs.AI ·
Foresight Without Seeing: Latent Futures for World Action Models
תקציר מקורי באנגליתarXiv:2608.11605v2 Announce Type: replace Abstract: World Action Models (WAMs) connect visual prediction with robot control, but supplying predictive context often requires expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. We introduce ForeWAM, a World Action Model that separates forecasting from rendering to expose and shape latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key-value states throughout action denoising. To make this context relevant to control, we introduce dynamics registers supervised by latent actions from a frozen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית