יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LC-QAT: תרגיל נתונים-קומפקטי 2-ביט QAT ל-LLMs דרך קונטרול-לינארי של וקטור קוויזציה

LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
LC-QAT היא תכונה נתונים-קומפקטי 2-ביט QAT ל-LLMs שמשתמשת בקונטרול-לינארי של וקטור קוויזציה. היא מספקת התאמה גבוהה של פקודות-קצבה ומאפשרת אופטימיזציה מלאה-לקוונטית. ניסויים שונים מדגימים את יעילותה של LC-QAT.
תקציר מקורי באנגליתarXiv:2606.10531v3 Announce Type: replace Abstract: Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on scalar quantization (SQ), which enables efficient optimization but suffers from severe performance degradation at 2-bit precision. On the other hand, vector quantization (VQ) provides substantially higher representational capacity, but its discrete codebook lookup prevents end-to-end training. We propose LC-QAT, a 2-bit weight-only VQ-QAT framework that represents quantized weights via a learned affine mapping over discrete vectors, which yields a high-quality PTQ initialization and enables fully differentiable end-to-end optimization without explicit codebook lookup in the training forward pass. This
קרא במקור המקורי