כתבה
arXiv cs.LG ·
Latent Goal Prediction from Language for Model-Based Planning
תקציר מקורי באנגליתarXiv:2606.20627v2 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) enable agents to plan in latent space by imagining the outcomes of candidate actions, yet task specification remains a bottleneck. Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on distinct large generative models. We introduce LAGO (Latent Goal Prediction from Language), a hierarchical world model in which a single predictor both forecasts action-conditioned dynamics and grounds language instructions as sequences of intermediate latent subgoals, training both modes with a single regression objective over a shared latent space. At each planning step, LAGO predicts a sequence
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית