כתבה
arXiv cs.AI ·
אופטימיזציה של דגם-עולם ישיר: למידה של העולם מעבר להתאימות לפעולה
Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
מחקר חדש הציע פרדיגמה ללמידה פוסט-שחרור של דגם-עולם, המשפרת את התאמת הפעולה לעולם. הפרדיגמה, שנקראת DEWO, משתמשת בניסיון חזותי כדי לשפר את הדגם-עולם, ולאפשר לו להתאים את הפעולה לעולם. המחקר הראה ש-DEWO משפרת את ההצלחה של הדגם-עולם במשימות שונות.
תקציר מקורי באנגליתarXiv:2609.37398v1 Announce Type: new Abstract: World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית