יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימון LLM א-סינכרוני

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
חוקרים מציגים שיטה חדשה לאימון LLM א-סינכרוני, GMC-GRPO, שמשפרת את העמידות לרולאוטים מיושנים. השיטה נבדקת על מודלים Qwen3 ומראה תוצאים טובים.
תקציר מקורי באנגליתarXiv:2610.01896v1 Announce Type: cross Abstract: Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-ba
קרא במקור המקורי