כתבה
arXiv cs.CL ·
שימור מידע מומחים ארוך-זנב בכיוונון Mixture-of-Experts
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
חוקרים הציגו שיטה חדשה לכיוונון מודלים Mixture-of-Experts, המשמרת מידע מומחים ארוך-זנב. השיטה משלבת ביאס-דרייבן ספארסיפיקציה עם גייטד קונדנסר אקספרטים. התוצאות מראות שיפור של 2.5%+ בבנצ'מרקים שונים.
תקציר מקורי באנגליתarXiv:2604.23036v2 Announce Type: replace-cross Abstract: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layers are fragile. Methods such as DenseMixer and ESFT mitigate router collapse with dense mixing or auxiliary load-balancing losses, but these introduce noisy gradients that often degrade performance. In preliminary experiments, we systematically pruned experts and observed that while certain super experts are activated far more frequently, discarding less used experts still leads to notable performance degradation. This suggests that even rarely activated experts encode non-trivial knowledge useful for downstream tasks. Motivated by this, we propose an auxiliary-loss-free MoE SFT framework that c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית