יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ניתוח התכנסות PPO עם מבקרים וקיטוע

A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
PPO-Clip הוא אלגוריתם RL נפוץ. מחקר זה מציג ניתוח התכנסות לא-אסימפטוטי של PPO-Clip. הניתוח כולל למידת מבקרים, קיטוע ושימוש חוזר בניסויים. התוצאות משפרות את ההבנה של PPO ומספקות הדרכה תאורטית לכיוון האלגוריתם.
תקציר מקורי באנגליתarXiv:2610.10273v1 Announce Type: new Abstract: Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood. We develop a \emph{non-asymptotic} analysis of PPO-Clip as a \emph{closed-loop actor--critic} system. It captures actor--critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets. Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters. A sufficient coupling
קרא במקור המקורי