יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Grounded World Model: Latent Planning with Language Goals

מודל עולם נעול נותן את האפשרות לתכנון חסר-ספק מבוסס-לשון בלבד. המודל, המבוסס על דגם-לשון-וידאו מוכנן מראש, מציג את העתיד במרחב הוויזואלי ומחשב את העלויות של כל פעולה. המודל נלמד רק מזוגות וידאו-פעולה, ולא צריך תווי-לשון. בניסויים סימולטוריים, המודל פתר 87% מ-288 משימות שלא נראו לפני כן, בעוד 10 דגמי-VLAs שהוטמעו פתרו 22%. המודל גם נבחן בסימולציה-איסאק ובסצנות אמיתיות.
תקציר מקורי באנגליתarXiv:2604.11751v2 Announce Type: replace-cross Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WI
קרא במקור המקורי