כתבה
arXiv cs.AI ·
Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence
תקציר מקורי באנגליתarXiv:2605.15618v2 Announce Type: replace-cross Abstract: Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reported as a final task score, obscuring how their representations behave under the degraded and ambiguous conditions a deployed world model must handle. We present the first systematic study of this component, analyzing four matched-capacity frontier self-supervised learning models that are strong candidates for world-model encoders -- V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2 -- across five robustness axes relevant to their deployment as video world models: feature discriminability, corruption robustness, fine-grained discrimination, occlusion robustness, and sensitivity to temporal directio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית