יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

תקציר מקורי באנגליתarXiv:2606.26472v4 Announce Type: replace Abstract: Reasoning models can generate chains of thought tens of thousands of tokens long, making the key--value (KV) cache that holds them a major bottleneck for inference throughput. Existing eviction policies for long reasoning traces typically rank cached tokens using attention weights, requiring access to the attention matrix and making them incompatible with fast inference kernels. In this work we study the limits of such policies under tight cache budgets. Surprisingly, we find that under the strongest of them the generations that finish are wrong about as often as without eviction; most of the accuracy loss comes from generations that enter loops and run until the length limit, and retaining more tokens according to a fixed importance scor
קרא במקור המקורי