כתבה
arXiv cs.CL ·
LILA: גזירה מובנית ללא כיוון של מודלי שפה גדולים
LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry
LILA היא שיטה חדשה לגזירה מובנית של מודלי שפה גדולים. היא משתמשת בגאומטריה ספקטרלית ומשיגה תוצאות טובות יותר משיטות אחרות, כמו PruneNet ו-SliceGPT, בלי צורך בנתוני כיוון. LILA עובדת עם מודלים כמו LLaMA-2-7B ו-Phi-2.
תקציר מקורי באנגליתarXiv:2609.11163v1 Announce Type: cross Abstract: Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp acr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית