כתבה
arXiv cs.AI ·
AutoLoCo: תקשורת יעילה באימון LLM בעזרת סינכרון אדפטיבי
AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
AutoLoCo מצליח להפחית את תדירות התקשורת באימון LLM על ידי אימוץ של פרקטיקה אדפטיבית.
תקציר מקורי באנגליתarXiv:2609.36662v1 Announce Type: cross Abstract: The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית