יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אימון LLM א-סינכרוני: ניתוח התכנסות

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
חוקרים הציגו שיטה חדשה לאימון LLM א-סינכרוני, GMC-GRPO, שמשפרת את התכנסות המודל. השיטה הודגמה על מודל Qwen3 והראתה עמידות טובה יותר לרולאוטים מיושנים.
תקציר מקורי באנגליתarXiv:2610.01896v1 Announce Type: new Abstract: Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-base
קרא במקור המקורי