יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בקרת יעילות עכירות ללמידת חיזוק א-סינכרונית

Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
PACE הוא אלגוריתם לבקרת עכירות יעילה בלמידת חיזוק א-סינכרונית. הוא משפר את הדיוק ב-18.7% בהשוואה לגישה א-סינכרונית רגילה. PACE תומך במודלים שונים, כולל LLM.
תקציר מקורי באנגליתarXiv:2609.36830v1 Announce Type: new Abstract: Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess po
קרא במקור המקורי