יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

EpiKV: פינוי קודם-ערך ללא מטריצת התנייה

EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix
EpiKV: פינוי קודם-ערך שמודע לאפיפניה ללא מטריצת התנייה. פיתוח חדש של פינוי קודם-ערך שמאפשר עבודה יעילה יותר של מודלי LLM.
תקציר מקורי באנגליתarXiv:2606.26472v4 Announce Type: replace-cross Abstract: Reasoning models can generate chains of thought tens of thousands of tokens long, making the key--value (KV) cache that holds them a major bottleneck for inference throughput. Existing eviction policies for long reasoning traces typically rank cached tokens using attention weights, requiring access to the attention matrix and making them incompatible with fast inference kernels. In this work we study the limits of such policies under tight cache budgets. Surprisingly, we find that under the strongest of them the generations that finish are wrong about as often as without eviction; most of the accuracy loss comes from generations that enter loops and run until the length limit, and retaining more tokens according to a fixed importanc
קרא במקור המקורי