יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SOAP, Muon ומעבר: דחיפת קנה מידה של LLM Pretraining

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
SOAP ו-Muon מציעים התכנסות מהירה יותר מ-AdamW, אך עלותם החישובית ואתגרי יציבות מספרית מוגבלים אימוץ בקנה מידה גדול. המחקר מציג שיפורים אלגוריתמיים לייצוב אימון בקנה מידה גדול.
תקציר מקורי באנגליתarXiv:2607.20548v1 Announce Type: new Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. In this work, we adapt and enhance preconditioned gradient methods to overcome the practical challenges of large-scale LLM pretraining. We first identify instabilities in SOAP at large batch sizes and propose algorithmic modifications including per-step QR orthogonalization and improved preconditioning strategies that eliminate loss spikes and enable stable training in these regimes. We then present a unified empirical study of SOAP, Muon, and AdamW using update-RMS matching to ensure fair learning rate transfer across optimizers. As part of this analysis, we empiric
קרא במקור המקורי