כתבה
arXiv cs.LG ·
עודף וסינרגיה: בחירת מומחים תלוית תלות
Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
DS-MoE הוא כלי חדש לבחירת מומחים במודלים MoE. הוא משתמש באופטימיזציה תת-מודולרית כדי למצוא את הקבוצה האופטימלית של מומחים. הכלי מצליח לשמור על יחסי גומלין חשובים בין המומחים ולהשיג ביצועים טובים יותר מאשר שיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.00558v1 Announce Type: new Abstract: While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redund
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית