כתבה
arXiv cs.LG ·
Communication-Efficient LLM Adaptation over Decentralized GPU Meshes
תקציר מקורי באנגליתarXiv:2609.14339v1 Announce Type: new Abstract: Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes the primary bottleneck. We study post-pretraining adaptation in this setting. We propose an asynchronous two-circuit system: a fast compressed training circuit drives throughput using activation masking for pipeline-parallel (PP) transfer and compressed data-parallel (DP) synchronization, while a slow anchor circuit runs occasional unmasked forward--backward passes off the critical path. Then, we introduce a spectral correction optimizer that uses these delayed anchor priors to denoise masked gradients without blocking the fast stream. Although prior work has found
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית