כתבה
arXiv cs.LG ·
RAZOR: Pruning Replaceable Experts in LLMs
תקציר מקורי באנגליתarXiv:2609.30465v3 Announce Type: replace Abstract: Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover. We introduce RAZOR, a training-free method that scores replaceability from consensus residuals, the deviations of individual expe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית