כתבה
arXiv cs.CL ·
RAZOR: גזירת חוקרים תחלופיים ב-LLMs
RAZOR: Pruning Replaceable Experts in LLMs
RAZOR - גזירת חוקרים תחלופיים ב-LLMs. מתודה זו מאפשרת גזירת חוקרים ב-LLMs תוך שמירה על יכולת המודל להסביר. המתודה, RAZOR, נבחנה במספר מודלים, כולל GLM-4.7-Flash ו-Qwen3.6-35B-A3B.
תקציר מקורי באנגליתarXiv:2609.30465v3 Announce Type: replace-cross Abstract: Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover. We introduce RAZOR, a training-free method that scores replaceability from consensus residuals, the deviations of individua
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית