יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קופלינג מומחה ב-MoE: צמצום עלות כל-לכל עם פלצורה מתאמה והזזת טוקן

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
קופלינג מומחה ב-MoE מצמצם את עלות הכל-לכל באמצעות פלצורה מתאמה והזזת טוקן. שיטה זו נבחנה ב-Megatron-LM.
תקציר מקורי באנגליתarXiv:2610.09372v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected toget
קרא במקור המקורי