כתבה
arXiv cs.LG ·
לכל 1-ביט: כלפי קיצור גישור 1-ביט אמיתי ל-LLMs
All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs
מאמר חדש מציג שיטה חדשה לקיצור גישור של LLMs, המצליחה להשיג קיצור 1-ביט עם חוסר פחות מ-1%.
תקציר מקורי באנגליתarXiv:2609.06161v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. Weight binarization is a promising solution, but existing binarization-based post-training quantization (PTQ) methods usually far exceed the nominal 1-bit storage target due to hidden overhead. To address this gap, we propose All for 1-Bit (AF1), a genuine 1-bit PTQ framework for LLMs. AF1 comprises two complementary components: (1) Null-space-Aware Binary Factorization (NABF) for improving binary reconstruction through Hessian-aware surrogate reparameterization, null-space-aware binary factorization, and scale-only global reconstruction; and (2) Hierarchical Shapley Allocation (HiSA) for as
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית