כתבה
arXiv cs.AI ·
הגברת פירוק של אקספרטים במודלי שפה של מערכת-אקספרטים
Higher-order pruning of experts in mixture-of-experts language models
במאמר זה פותחה שיטה חדשה להגברת פירוק של אקספרטים במודלי שפה של מערכת-אקספרטים, כדי להפחית את כמות הפרמטרים. השיטה, הנקראת HOPE, מצליחה לשמור על יעילות גבוהה גם בפירוק גבוה.
תקציר מקורי באנגליתarXiv:2609.18916v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית