כתבה
arXiv cs.AI ·
KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
תקציר מקורי באנגליתarXiv:2609.37988v2 Announce Type: replace Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית