כתבה
arXiv cs.AI ·
Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
תקציר מקורי באנגליתarXiv:2609.37751v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additiona
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית