כתבה
arXiv cs.LG ·
שימוש מחדש בדגימות קודמות באופטימיזציה פרוקסימלית: כאשר ואיך הוא עוזר?
Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?
במאמר זה, נחקרת יעילות שימוש מחדש בדגימות קודמות באופטימיזציה פרוקסימלית (PPO) והתוצאות המשמעותיות שלה. המחקר מציג שני גרסאות של PPO שמשתמשות בשימוש מחדש בדגימות ומציעות תאוריה תומכת. התוצאות המדעיות המצוינות במאמר זה תומכות בכך ששימוש מחדש בדגימות יעילות ומשפרת את התוצאות הסופיות של PPO.
תקציר מקורי באנגליתarXiv:2610.01399v1 Announce Type: new Abstract: Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic stu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית