כתבה
arXiv cs.LG ·
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
תקציר מקורי באנגליתarXiv:2511.10333v2 Announce Type: replace Abstract: Training large language models (LLMs) at scale incurs substantial communication overhead, while static gradient compression cannot adapt to gradient evolution and may degrade model quality. We propose EDGC, an entropy-driven dynamic gradient compression framework that adapts compression ranks to gradient entropy during training. EDGC combines efficient entropy estimation through gradient sampling, a theoretical model relating entropy to compression rank under a bounded-error constraint, and window-based rank adjustment across pipeline stages. Experiments on 32-V100 and 64-H100 GPU clusters training GPT2 models with 2.5B and 12.1B parameters show that EDGC reduces communication latency by up to 46.45% and end-to-end training time by 16.13%
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית