כתבה
arXiv cs.LG ·
MoE צפוף: חיפוש גרנולריות בקומפיוט קבוע
Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute
מודלי MoE צפופים הציגו אבדן התאמה לא-מונוטוני עם גודל המומחים
תקציר מקורי באנגליתarXiv:2610.02584v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnosti
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית