כתבה
arXiv cs.LG ·
WorldPack: Dynamic Frame Compression for Long-context Video World Modeling
תקציר מקורי באנגליתarXiv:2512.02473v2 Announce Type: replace-cross Abstract: Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. However, achieving temporally and spatially consistent generation over long horizons remains an open challenge: existing approaches either compress past frames using generic importance schedules that do not exploit 3D viewpoint geometry, or retrieve only a handful of spatially relevant frames without increasing the total amount of retained history. In this paper, we propose WorldPack, a video world model that introduces spatially-aware compressed memory to address both limitations simultaneously. The key insight is that compression rates should not be u
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית