כתבה
arXiv cs.CL ·
World Embedding Benchmark
תקציר מקורי באנגליתarXiv:2610.03632v1 Announce Type: cross Abstract: Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית