יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

MoRE: הגדלת מיקסר של מומחים עם ניתוב בעל דרגה נמוכה

MoRE: Scaling mixture of experts with hardware-aware low-rank routing
MoRE הוא אלגוריתם חדש שמשפר את ביצועי מודלי שפה על ידי ניתוב בעל דרגה נמוכה. הוא מאפשר הגדלת מספר המומחים במודל, מבלי לפגוע בביצועים. האלגוריתם נבדק על משימות שונות והראה שיפורים משמעותיים.
תקציר מקורי באנגליתarXiv:2609.36301v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a s
קרא במקור המקורי