כתבה
arXiv cs.LG ·
How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms
תקציר מקורי באנגליתarXiv:2606.09646v2 Announce Type: replace-cross Abstract: We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית