יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

RAZOR: פרוסת מומחים תחלופיים ב-LLMs

RAZOR: Pruning Replaceable Experts in LLMs
RAZOR היא שיטת פרוסת מומחים תחלופיים ב-LLMs שמשמשת קונצנסוס של תוצאות מומחים. היא מציגה תוצאות טובות יותר משיטות פרוסה אחרות, כולל REAP.
תקציר מקורי באנגליתarXiv:2609.30465v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refil
קרא במקור המקורי