כתבה
arXiv cs.LG ·
Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training
תקציר מקורי באנגליתarXiv:2610.09183v1 Announce Type: new Abstract: Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraini
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית