כתבה
arXiv cs.LG ·
Faster-WAM: Do World Action Models Need Deep Action Modules?
תקציר מקורי באנגליתarXiv:2608.02365v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית