כתבה
arXiv cs.CL ·
RAZOR: גזירת מומחים תחלופיים ב-LLMs
RAZOR: Pruning Replaceable Experts in LLMs
RAZOR היא שיטת גזירה ל-LLMs המבוססת על תזוזות רצפיות, המתארות את השינויים בתוצאות המומחים. השיטה מאפשרת גזירה של 25% ו-50% מהמומחים בלי לפגוע בביצועי המודל.
תקציר מקורי באנגליתarXiv:2609.30465v4 Announce Type: replace-cross Abstract: Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replacement experts compensate for the removed output. We introduce RAZOR, a training-free pruning method based on consensus residuals, the deviations of expert outputs from their original weighted mixture. At a fixed layer input, these residuals give the exact output change for a single deletion under survivor renormalization and router refill. RAZOR aggregates this damage by conditional root mean square and selects experts
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית