יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

LILA: צמצום ספקטרלי ללא-קליברציה של דגימות שפה גדולות

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry
LILA מציגה צמצום ספקטרלי ללא-קליברציה של דגימות שפה גדולות, כולל שימוש ב-LLaMA.
תקציר מקורי באנגליתarXiv:2609.11163v1 Announce Type: new Abstract: Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp acros
קרא במקור המקורי