יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

איך מודלים יסודיים של וידאו מקודדים פיזיקה אינטואיטיבית?

How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms
חוקרים את האופן שבו מודלים יסודיים של וידאו מקודדים מידע על פיזיקה אינטואיטיבית. המחקר בודק את הביצועים של מודלים שונים, כולל V-JEPA, VideoMAE ו-LTX-Video, ומצא ש-V-JEPA משיג את התוצאות הטובות ביותר.
תקציר מקורי באנגליתarXiv:2606.09646v2 Announce Type: replace-cross Abstract: We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible a
קרא במקור המקורי