כתבה
arXiv cs.LG ·
JARQ: שיפור עידון משותף לקוונטיזציה
JARQ: Joint Alternating Refinement for Quantization
JARQ הוא אלגוריתם לשיפור קוונטיזציה של מודלי שפה גדולים. הוא משפר את דיוק המודלים Llama-2, Llama-3 ו-Qwen. JARQ מוריד את ה-perplexity ב-90 מתוך 96 השוואות ומשפר את הדיוק של מודלים אלו.
תקציר מקורי באנגליתarXiv:2609.38599v2 Announce Type: replace Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית