כתבה
arXiv cs.LG ·
קונסולידציה של מדיניות רציפה ללמידה חיים-ארוכה של רובוט
Continual Policy Consolidation for Lifelong Robot Learning
אנו מציגים פרקטיקה חדשה לקונסולידציה של מדיניות רציפה ללמידה חיים-ארוכה של רובוט. הפרקטיקה, CPC, משלבת חידושים של רפורמינג למידה ושימור ידע. CPC נבחן באמצעות תצוגה של רובוטים ומציג תוצאות טובות. CPC יכול לשמש ככלי לקונסולידציה של מדיניות רציפה ללמידה חיים-ארוכה של רובוט.
תקציר מקורי באנגליתarXiv:2601.22475v2 Announce Type: replace Abstract: Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing a single policy over a growing task stream is difficult because robotic interaction is expensive, task distributions are heterogeneous, and sequential updates induce interference. To address these problems, we propose continual policy consolidation (CPC), a teacher--student framework that combines continual policy distillation with prioritized experience replay and expandable experts. This architecture separates skill acquisition from policy consolidation: teachers are trained independently through reinforcement learning, and their behaviors are continually distilled into a central generalist
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית