יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

MoRA: גזימת MoE דרך למידת הטיה של ראוטר וקירוב מומחים

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation
MoRA היא שיטה חדשה לגזימת מומחים במודלים MoE. היא משתמשת בלמידת הטיה של ראוטר וקירוב מומחים כדי לשפר ביצועים. MoRA נבדקה על מודלים Qwen3-30B-A3B, DeepSeek-V2-Lite ו-Moonlight-16B-A3B.
תקציר מקורי באנגליתarXiv:2610.00367v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expe
קרא במקור המקורי