יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הקטנת חד-פעמית של חוקרים מומרים במודלי Mixture-of-Experts דקני

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
במאמר זה, המחברים מציגים שיטה להקטנת חד-פעמית של חוקרים מומרים במודלי Mixture-of-Experts דקני. השיטה, הנקראת 'Training-Free Halving', מאפשרת הקטנת חד-פעמית של חוקרים מומרים במודלי Mixture-of-Experts דקני, תוך שמירה על תפקודם. המחברים מציגים תוצאות ניסויים שמעידות על יעילות של השיטה. המאמר עוסק בפיתוח של מודלי Mixture-of-Experts דקני ובשיטות להקטנת חד-פעמית של חוקרים מומרים במודלי Mixture-of-Experts דקני.
תקציר מקורי באנגליתarXiv:2609.04575v1 Announce Type: new Abstract: Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities. We show that this renormalization implicitly calibrates expert output gain to the training top-$k$: reducing $k$ at inference changes not only which experts are used but also the strength of the expert branch. We separate these effects by activating the top $k_1$ experts while normalizing by the probability mass of the top $k_2$ experts, introducing one integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, reducing from 8 to 4 experts causes a 4.65-point MMLU drop under standard renormalization but only 0.35 points with $k_2=16$, while halving routed-expert compute. T
קרא במקור המקורי