כתבה
arXiv cs.AI ·
Tailoring the Quantization Space for 1-Bit KV Cache Compression
תקציר מקורי באנגליתarXiv:2610.03027v1 Announce Type: cross Abstract: The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית