כתבה
arXiv cs.AI ·
כאשר קטן רציון השיקום פוגע: רפינר תפוצה רב-גונית להקטנת תצורה נמוכה של LLM
When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
נמצא כי רציון שיקום נמוך יכול לפגוע בביצועי המודל. נציגה חדשה, DRQ, משפרת את קודי המספור של המשקלים המוקטנים כדי לשפר את הביצועים הסופיים.
תקציר מקורי באנגליתarXiv:2610.11226v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית