כתבה
arXiv cs.LG ·
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
תקציר מקורי באנגליתarXiv:2607.24377v1 Announce Type: new Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quant
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית