יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

MoRE: פיתוח של מערכת של אקספרטים עם רוטינג נמוך-דרגה

MoRE: Scaling mixture of experts with hardware-aware low-rank routing
MoRE (Mixture of Rank-reduced-routed Experts) - פיתוח של מערכת של אקספרטים שמשתמשת ברוטינג נמוך-דרגה. המערכת מפחיתה את עלויות המחשוב ומאפשרת פיתוח של מודלים גדולים יותר. המחקר נערך על ידי Microsoft AI ומשתמש ב-GPU של NVIDIA.
תקציר מקורי באנגליתarXiv:2609.36301v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a syn
קרא במקור המקורי