כתבה
arXiv cs.AI ·
WAM-OPD: ניהול פוסט-אימון של דגם עצומי-פעולה עם סיכוך על-מדיני
WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation
WAM-OPD מציג ניהול פוסט-אימון של דגם עצומי-פעולה, עם סיכוך על-מדיני. המאמר מציג תוצאות של 65.7% במשימות RoboTwin 2.0 ו-64.6% במשימות רובוטים אמיתיים.
תקציר מקורי באנגליתarXiv:2608.22364v2 Announce Type: replace Abstract: World Action Models (WAMs) generate both future video and robot actions, offering two connected outputs for post-training supervision. How can a pretrained WAM learn from a stronger Teacher on the histories it encounters during execution? We present WAM-OPD, which collects Student rollout histories and queries a Teacher for paired video and action targets. The Student learns from both targets while retaining its one-step video and action generation at deployment. Across 12 RoboTwin 2.0 tasks, WAM-OPD improves average success from 33.8% to 65.7%; across four real-robot tasks, it improves average success from 51.4% to 64.6%. With the collected Student histories held fixed, joint video-action supervision achieves the highest observed success
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית