כתבה
arXiv cs.LG ·
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
תקציר מקורי באנגליתarXiv:2605.06206v2 Announce Type: replace Abstract: Mixture of experts has emerged as the primary mechanism for making Large Language Models (LLMs) computationally efficient. However, in distributed settings, communicating token embeddings between experts is a significant bottleneck. We present the novel Federation of Experts (FoE) architecture. FoE restructures the MoE block of a transformer layer into multiple MoE clusters. Each cluster is responsible for only one of the KV heads and expert parallelism is applied between those experts. Between clusters, a sum synchronizes the post-attention residuals, which then drives routing and dispatch for the next MoE block. In a single-node setting, FoE completely eliminates all-to-all communication as all experts within a group are contained on th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית