כתבה
arXiv cs.LG ·
התאמת המרחב של הקיצוניות להקלה של 1-ביט KV Cache
Tailoring the Quantization Space for 1-Bit KV Cache Compression
אנו מציגים את TaSQ, שמתאים את המרחב של הקיצוניות להקלה של 1-ביט KV Cache. השיטה שלנו משלבת עריכה מודולרית של קנבוק, תכנון מחדש של קנבוק, והקבצה של קנבוק. TaSQ נותן תוצאות טובות יותר מאשר השיטות הקיימות להקלה של KV Cache.
תקציר מקורי באנגליתarXiv:2610.03027v1 Announce Type: new Abstract: The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית