כתבה
arXiv cs.LG ·
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
תקציר מקורי באנגליתarXiv:2610.11315v1 Announce Type: new Abstract: In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית