כתבה
arXiv cs.CL ·
הפשטת גבולות החשיפה העליונה של Top-k למיקסטור
Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
אנו מציגים את Elastic Expert Routing, שמטרתו להפשיט את גבולות החשיפה העליונה של Top-k למיקסטור. זה עוזר לשפר את הביצועים של מודלי Mixture-of-Experts.
תקציר מקורי באנגליתarXiv:2610.11575v1 Announce Type: new Abstract: Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distributi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית