כתבה
arXiv cs.AI ·
קיצוץ מומחים במודלים MoE
When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
חוקרים גילו כי קיצוץ מומחים במודלים Mixture-of-Experts (MoE) יכול להיות בעייתי כאשר הרוטר מפצל את הנתונים באופן אחיד מדי. הם הציעו שיטה חדשה לקיצוץ, Minimax Expert Score Allocation (MESA), שמשפרת את הדיוק במגוון משימות.
תקציר מקורי באנגליתarXiv:2609.04453v1 Announce Type: cross Abstract: Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities provide a reliable importance signal. We observe that this assumption breaks down under over-dispersed routing, a regime associated with aggressive load-balancing during training, in which tokens are distributed nearly uniformly across experts and importance signals collapse. In this regime, perplexity does not predict downstream task accuracy: on gpt-oss-20B, the lowest-perplexity pruning configuration yields the worst mathematical reasoning, while the highest-perplexity configuration preserves it. This does not occur under standard routing (e.g., Mixtral-8x7B-Instr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית