כתבה
arXiv cs.LG ·
How Long, Not How Close: A Learned Temporal Metric for Planning in Latent World Models
תקציר מקורי באנגליתarXiv:2610.04988v2 Announce Type: replace Abstract: Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the recorded trajectories already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית