כתבה
arXiv cs.AI ·
כל ל-1: בדרך לקומפרציה 1-ביט נכונה ל-LLMs
All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs
אנו מציגים פרקטיקה 1-ביט נכונה לקומפרציה של LLMs, המשמרת דיוק המודלים תחת 1.0-BPW.
תקציר מקורי באנגליתarXiv:2609.06161v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. Weight binarization is a promising solution, but existing binarization-based post-training quantization (PTQ) methods usually far exceed the nominal 1-bit storage target due to hidden overhead. To address this gap, we propose All for 1-Bit (AF1), a genuine 1-bit PTQ framework for LLMs. AF1 comprises two complementary components: (1) Null-space-Aware Binary Factorization (NABF) for improving binary reconstruction through Hessian-aware surrogate reparameterization, null-space-aware binary factorization, and scale-only global reconstruction; and (2) Hierarchical Shapley Allocation (HiSA) for
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית