כתבה
arXiv cs.AI ·
דגימת עולם-אגו לייצור וידאו מגוף למשימות ניווט-טיפול באורך-זמן
World-Ego Modeling for Embodied Video Generation in Long-Horizon Navigation-Manipulation Tasks
ה-WEM (World-Ego Model) הוא דגימת עולם-אגו שמטפלת בייצור וידאו מגוף למשימות ניווט-טיפול באורך-זמן. ה-WEM משתמש במודלי RCA ו-SR-MoE כדי לפרד בין העולם והאגו, ומציג תוצאות טובות ב-HTEWorld.
תקציר מקורי באנגליתarXiv:2605.19957v2 Announce Type: replace-cross Abstract: Embodied video world models typically capture both scene evolution and the robot's behavior, which we refer to as the \emph{world} and the \emph{ego}, respectively. The world and the ego exhibit different underlying dynamics: world prediction relies primarily on visual history and emphasizes scene stability, whereas ego prediction relies more strongly on the current instruction and emphasizes accurate instruction following. Modeling both components within a single generation stream can entangle these different dependencies, making it difficult to specialize the prediction of either component. Consequently, it becomes difficult to simultaneously maintain scene consistency and accurate instruction following, particularly in long-horiz
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית