יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דחיסת MoE ללא אימון מחדש: סקירה מושווית

Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment
חוקרים בדקו דחיסת MoE ללא אימון מחדש, ומצאו כי הוספת שלב התאמה קטן לאחר דחיסה משפר את הביצועים. הם השוו בין שיטות שונות, כולל אימון מחדש והעברת ידע, ומצאו כי התאמה מלאה של פרמטרים נותנת את התוצאות הטובות ביותר.
תקציר מקורי באנגליתarXiv:2609.06076v1 Announce Type: new Abstract: Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a tiny post-compression adjustment stage. Across two MoE LLM backbones, four pruning/merging methods, three expert-retention ratios, and 28 benchmarks, we compare LM fine-tuning and teacher-based KD under matched small-data budgets and measured GPU costs. Using only 3,000 C4 examples and a single epoch of adjustment, Full FT recovers 37.3% of the original-to-compressed performance gap on average. Moreover, LM fine-tuning is more cost-effective than standard
קרא במקור המקורי