כתבה
arXiv cs.LG ·
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
תקציר מקורי באנגליתarXiv:2609.00049v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q ta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית