כתבה
arXiv cs.CL ·
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
תקציר מקורי באנגליתarXiv:2510.02345v4 Announce Type: replace Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on dynamic expert clustering and structured compression to address these issues cohesively. Our method employs an online clustering procedure that periodically regroups experts using a fused metric of parameter and activation similarity, which stabilizes expert utilization. To our knowledge, this is one of the first frameworks to leverage the semantic embedding capability of the router to dynamically reconfigure the model's architecture during training for substantial efficiency gains. Within each cluster, we decompose expert weights into a shared base matrix and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית