יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תכנון עולם תלוי-לשון לניווט ראייתי

Language-Conditioned World Modeling for Visual Navigation
חידוש זה עוסק בניווט ראייתי תלוי-לשון לאגנטים גופניים. המחקר מציג דגם עולם תלוי-לשון (LCVN-WM) שמדמה צפיות עתידיות ואגנט-מבקר (LCVN-AC) שלומד את מדדיו כולם במרחב הלטנטי המדומה. המחקר גם מציג פרדיקטיבה אוטורג'סיבית (LCVN-Uni) שמציע פורטרפוליאו של פעולות וצפיות באופן אחד בעבור תפריט תקין.
תקציר מקורי באנגליתarXiv:2603.26741v2 Announce Type: replace-cross Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control. We introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human-verified instructions spanning diverse environments and instruction styles. Building on this benchmark, we study two complementary paradigms: (i) latent-imagination policy learning, in which a diffusion-based world model (LCVN-WM) imagines future obs
קרא במקור המקורי