כתבה
arXiv cs.CL ·
הצפנת LLM על ידי הסרת רשתות עם אופטימיזציה בינארית מוגבלת
LLM Compression by Block Removal with Constrained Binary Optimization
במאמר זה, פותחים טכניקה להצפנת LLM על ידי הסרת רשתות, ומציגים תוצאות טובות בהשוואה לטכניקות אחרות. הטכניקה נקראת LLM Compression by Block Removal with Constrained Binary Optimization.
תקציר מקורי באנגליתarXiv:2602.00161v3 Announce Type: replace-cross Abstract: In this paper, we formulate the compression of large language models (LLMs) by optimally deleting transformer blocks (``block removal'') as a constrained binary optimization (CBO) problem that can be mapped to a physical system (Ising glass), whose energies are a strong proxy for downstream model performance. This formulation enables an efficient ranking of a large number of candidate block-removal configurations yielding many high-quality, non-trivial solutions beyond those only removing consecutive regions. Our method performs strongly in the deep compression regime, such as for 50% compression of Llama-3.3-70B-Instruct, where we achieve an almost 23 percentage point increase on the MMLU benchmark compared to other state-of-the-ar
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית