כתבה
arXiv cs.CL ·
It Takes a MAESTRO To Prune Bad Experts
תקציר מקורי באנגליתarXiv:2607.08601v2 Announce Type: replace Abstract: Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework designed for MoE architectures that models autoregressive expert activation trajectories as Ergodic Markov chains whose stationary distributions encode cr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית