כתבה
arXiv cs.AI ·
RoboAware: למידה לקואורדינטירציה של מיומנויות גופניות מתוצאות נגדיות
RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes
RoboAware מלמד לקואורדינטירציה של מיומנויות גופניות מתוצאות נגדיות. המאמר מציג את RoboAware, שבנה על סקיל אורכסטרציה של סקילים רובוטיים, ומציע Execution-Aware Learning (EAL), שמשלב Q-learning עם Monte Carlo tree search. RoboAware הציג תוצאות טובות בבדיקות יחיד-פרק על 100 תפקידים, והציג תוצאות SOTA ב-RoboSuite, LIBERO-Pro ו-RoboTwin.
תקציר מקורי באנגליתarXiv:2610.11480v1 Announce Type: cross Abstract: Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית