יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חינוך חופשי, נכון על עצים: תיקון הפגם של PPO קונה כושר ניסויי תחת שימוש תוקפני מחדש

Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse
תיקון הפגם של PPO קונה כושר ניסויי תחת שימוש תוקפני מחדש. המחקר מציע תיקון חדש לפגם של PPO, שמקנה כושר ניסויי תחת שימוש תוקפני מחדש. התיקון נבחן במספר תרגילים קשים ומציג תוצאות משיכות.
תקציר מקורי באנגליתarXiv:2609.39634v1 Announce Type: cross Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autore
קרא במקור המקורי