כתבה
arXiv cs.LG ·
חינוך חופשי, נקי על עצים: תיקון הפגיעה של PPO קונה כושר ניסיוני תחת שימוש אגרסיבי
Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse
מחקר חדש מציע תיקון לאלגוריתם PPO, שמשפר את כושר הניסיון שלו תחת שימוש אגרסיבי. התיקון נועד לפגוע בטעות שנוצרת כאשר האלגוריתם משתמש בנתוני ניסויים רחוקי-מדורג. המחקר מציע תיקון חדש, שמקטין את הטעות ומשפר את כושר הניסיון של PPO.
תקציר מקורי באנגליתarXiv:2609.39634v1 Announce Type: new Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית