כתבה
arXiv cs.AI ·
ASH: סוכנים שמשפרים את עצמם בעולמות של זמן ארוך
ASH: Agents that Self-Hone in Long-Horizon Worlds
ASH היא סוכנית שלומדת מדלגת-זמן-ארוך מסרטי וידאו לא מאותויים, ללא עיצוב שכר או הערכה מומחים. ASH משפרת את עצמה בעזרת דגם תגובה הפוך (IDM) שהיא לומדת מנתיביה שלה, ומשתמשת ב-IDM שלה כדי לגרוף הדרכה מסרטי וידאו רלוונטיים. ASH משתמשת בלמידה בלי הדרכה כדי לזהות רגעים חשובים מסרטי וידאו גדולי-היקף ולשמור אותם כזיכרון-ארוך - מאפשר לה להתמודד עם בעיות-זמן-ארוך.
תקציר מקורי באנגליתarXiv:2605.14211v4 Announce Type: replace Abstract: Long-horizon visuomotor tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We introduce ASH, an agentic system that learns a long-horizon policy from unlabeled, noisy internet video, without reward shaping or expert annotation. ASH follows a self-improvement loop; when it gets stuck, ASH learns an Inverse Dynamics Model (IDM) from its own trajectories, and uses its IDM to extract supervision from relevant internet video. ASH uses unsupervised learning to identify key moments from large-scale internet video and retains them as long-term memory - allowing it to tackle long-horizon problems. We evaluate ASH on two complementary environments
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית