יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ASH: Agents שמטפלים בעצמם בעולמות עם תקופות ארוכות

ASH: Agents that Self-Hone in Long-Horizon Worlds
ASH לומד פוליצי ארוך-תקופה מווידאו אינטרנט לא מאותויים. הוא משתמש בשיפור עצמי כדי ללמוד מסלולי תנועה ולקבל הנחיות מווידאו אינטרנט. ASH יכול לטפל בבעיות עם תקופות ארוכות.
תקציר מקורי באנגליתarXiv:2605.14211v5 Announce Type: replace Abstract: Long-horizon visuomotor tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We introduce ASH, an agentic system that learns a long-horizon policy from unlabeled, noisy internet video, without reward shaping or expert annotation. ASH follows a self-improvement loop; when it gets stuck, ASH learns an Inverse Dynamics Model (IDM) from its own trajectories, and uses its IDM to extract supervision from relevant internet video. ASH uses unsupervised learning to identify key moments from large-scale internet video and retains them as long-term memory - allowing it to tackle long-horizon problems. We evaluate ASH on two complementary environments
קרא במקור המקורי