יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הקאש של KV הוא הקיר החדש של המסך

The KV Cache Is the New Memory Wall
הקאש של KV מגביל את האינפרנסה של LLM באורכי תווים ארוכים. המחקר חוקר את הקשר בין קישור-ערך לזיכרון ומציע דרכים לשפר את הביצועים.
תקציר מקורי באנגליתarXiv:2609.30854v1 Announce Type: cross Abstract: Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128k-token sequence adds 42 GB of KV cache. Techniques that compress, evict, page, share, or offload KV state have proliferated, but reported gains use inconsistent workloads, hardware, and quality metrics, preventing cross-paper comparison. This SoK paper unifies the field analytically, with a protocol that strictly separates derived and reported claims. We derive closed-form arithmetic intensity as a decaying function of contex
קרא במקור המקורי