כתבה
arXiv cs.AI ·
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
תקציר מקורי באנגליתarXiv:2609.05764v1 Announce Type: cross Abstract: The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound. Holding a quantized KV cache in dense on-chip non-volatile memory (NVM) removes the off-chip transfer. Existing KV quantization methods, however, were designed for GPU-style memory systems: KIVI attaches per-group metadata, adding about 25% to the stored KV cache; KVQuant keeps sparse full-precision outliers that a dense array cannot hold in place. This paper examines what these structures cost when the KV cache resides in NVM behind fixed-range converters, and designs a quantization scheme matched to that interface. The architecture stores the quantized KV cache
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית