יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

G$^2$PTQ: שיפור קוונטיזציה של LLM לאחר הכשרה

G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
G$^2$PTQ מציע שיפור בקוונטיזציה של LLM לאחר הכשרה, על ידי שילוב של מידע ראשון ושני באופן גלובלי. המאמר מציג פרקטיקה חדשה של G$^2$PTQ, שמאפשרת שיפור באיכות הקוונטיזציה וביצועים טובים יותר. הקוד זמין בגיטהב: https://github.com/G2PTQ/G2PTQ.
תקציר מקורי באנגליתarXiv:2609.31009v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian es
קרא במקור המקורי