כתבה
arXiv cs.LG ·
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
תקציר מקורי באנגליתarXiv:2608.13966v2 Announce Type: replace Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT has a structural mismatch: gradient updates are applied to latent full-precision weights, while the loss and gradients are computed on lossy reconstructions of those weights. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods address a similar problem by minimizing loss-aware reconstruction error, but applying such expensive reconstruction repeatedly during QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that brings lightweight, loss-aw
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית