כתבה
arXiv cs.LG ·
RAZOR: עיוות ספקים תחלופיים ב-LLMs
RAZOR: Pruning Replaceable Experts in LLMs
RAZOR היא שיטת עיוות ספקים תחלופיים ל-LLMs, המבוססת על תלותיות ספקים. השיטה מציעה דרך עיוות ספקים תחלופיים ל-LLMs, כדי לשפר את יעילות המודל. השיטה נבחנה במודלים שונים, כולל GLM-4.7-Flash ו-Qwen3.6-35B-A3B.
תקציר מקורי באנגליתarXiv:2609.30465v4 Announce Type: replace Abstract: Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replacement experts compensate for the removed output. We introduce RAZOR, a training-free pruning method based on consensus residuals, the deviations of expert outputs from their original weighted mixture. At a fixed layer input, these residuals give the exact output change for a single deletion under survivor renormalization and router refill. RAZOR aggregates this damage by conditional root mean square and selects experts under
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית