יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פירוק שותפי של נקודות קצבה נמוכות-דרגה להקטנת חידושי MoE

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
אפשרות להקטנת MoE ללא נתונים על ידי פירוק שותפי של נקודות קצבה נמוכות-דרגה. השיטה מאפשרת שיתוף ריצוף רב-מומנטלי, התכווצות מהירה וטוהר נמוך יותר. השיטה החדשה, SLBF, נבחנה על חמש חידושי MoE שונות, והציגה תוצאות טובות יותר משיטות הקטנה של MoE שקדמו לה.
תקציר מקורי באנגליתarXiv:2610.09342v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence
קרא במקור המקורי