כתבה
arXiv cs.AI ·
iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
תקציר מקורי באנגליתarXiv:2610.02815v1 Announce Type: new Abstract: Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית