כתבה
arXiv cs.LG ·
שרשרת אסטרטגיות מקבילות להתכנסות אימון מהירה
Parallelism Strategy Chaining for Fast Training Convergence
CONA היא שיטה חדשה לאימון מודלים שבוחרת אסטרטגיות מקבילות בזמן אמת. היא מאפשרת התכנסות מהירה יותר של מודלים כמו GPT-3 ו-Llama. CONA משתמשת בנתוני גרדיאנט וקצב חישוב כדי לבחור את האסטרטגיה הטובה ביותר.
תקציר מקורי באנגליתarXiv:2609.07236v1 Announce Type: new Abstract: Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large language models. State-of-the-art methods search for a parallelism strategy offline and select the single strategy that minimizes per-iteration time. But we find that they neglect the target validation perplexity and time-to-perplexity (TTP). In particular, our analysis reveals that the best strategy yielding the fastest perplexity improvement changes multiple times during training. As a result, state-of-the-art methods are 1.8-11.4x slower in TTP than the strategy sequence that selects the best strategy at each iteration. This paper proposes CO
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית