כתבה
arXiv cs.LG ·
Action Chunking Proximal Policy Optimization with Feedback Correction
הצגנו פרוטוקול פופולרי (PPO) המשתלב עם חיתוך פעולות ותיקון נקלה של פעולות בתוך חיתוך. הפרוטוקול, ACPPO-Corr, משפר את הביצועים במשימות רובוטיות סימולטוריות.
תקציר מקורי באנגליתarXiv:2609.36250v1 Announce Type: new Abstract: Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts pl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית