כתבה
arXiv cs.LG ·
SoloQ: Calibration-Free Quantization for Diffusion Language Models
תקציר מקורי באנגליתarXiv:2610.07121v1 Announce Type: new Abstract: Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית