כתבה
arXiv cs.AI ·
One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
תקציר מקורי באנגליתarXiv:2609.14833v1 Announce Type: new Abstract: World models, systems that generate what happens next given current environmental conditions, are increasingly being implemented with multi-modal generation in mind. However, generating multiple modalities simultaneously, such as visual simulations alongside physical state predictions in the form of text, introduces the risk of cross-modal inconsistency. Tested separately, both outputs may look convincing while still disagreeing: a model can calculate that a ball should rebound in one modality, then generate no rebound in another modality, to say nothing of diverging from real-world dynamics entirely. In this work we focus on two failures explicitly: \emph{Internal misalignment}, the disagreement between the world model's generated video and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית