יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חסרות לוגיטים ישנים בRL אסינכרוני: תאושטות סמנטית ושיטות תיקון לתיקון-לא-פוליצי

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
במאמר זה נחקרה תאושטות סמנטית בRL אסינכרוני, ונציגו שיטות תיקון לתיקון-לא-פוליצי. נבחנו שלושה דרכי תיקון נקיות: עקיבה-בסנאפשוט, דגל-לוגיט-ישן, והפסקה-בפארט-רולאוט. נמצא כי שיטת PPO-EWMA, שהוצגה במאמר, משפרת את המהירות הלימוד ואת הביצועים של האופטימיזציה.
תקציר מקורי באנגליתarXiv:2605.12070v3 Announce Type: replace-cross Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emph{training--inference discrepancy term} that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emph{policy-staleness term} that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required histori
קרא במקור המקורי