כתבה
arXiv cs.LG ·
למידה לצורך תכנון מתוך חקירה סוררת
Learning to Plan from Random Exploration
במאמר זה, נראה כיצד ניתן ללמוד לתכנן מתוך חקירה סוררת, ללא צורך באימון תכנון-שיפור. המאמר כולל ניתוח של יחסי זמן, והצגה של דגם תנאי-אנרגיה רציונלי שמוערך על-ידי המודל. המאמר כולל גם תיאור של תכנון-מסלול ותכנון-גישה, והצגה של תוצאות ניסויים.
תקציר מקורי באנגליתarXiv:2609.38383v1 Announce Type: new Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the go
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית