כתבה
arXiv cs.AI ·
ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
תקציר מקורי באנגליתarXiv:2609.36835v1 Announce Type: new Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we fir
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית