יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מכשירים פרה-טריינד למיומנות: RL לפעילות רציפה עם מינימום התערבות אנושית

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
פיתוח של רשת רפרן עם RL לפעילות רציפה עם מינימום התערבות אנושית. המאמר עוסק בפיתוח של רשת רפרן שמסוגלת לבצע פעילות רציפה עם מינימום התערבות אנושית. המחברים מציגים את PARTS, פלטפורמה של RL שמסוגלת ללמוד מפעילות רציפה עם מינימום התערבות אנושית.
תקציר מקורי באנגליתarXiv:2609.21788v2 Announce Type: replace-cross Abstract: A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout e
קרא במקור המקורי