יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

טאילורינג המרחב הקוונטיזציה להקלה אגרסיבית של KV Cache

Tailoring the Quantization Space for 1-Bit KV Cache Compression
אנו מציגים את TaSQ, שהוא פרויקט של קוונטיזציה שמטפל בבעיית ההקלה האגרסיבית של KV Cache. TaSQ משלב עריכה מודולרית של קנבוק, תכנון מחדש של רשת, והקלה של קודבוק. הפרויקט נועד לשפר את יעילות הקוונטיזציה ולהקל את ההקלה של KV Cache.
תקציר מקורי באנגליתarXiv:2610.03027v1 Announce Type: cross Abstract: The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error
קרא במקור המקורי