יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SlimKV: דחיסת KV-קאש משותפת

SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention
SlimKV הוא שיטה חדשה לדחיסת KV-קאש, המאפשרת לשרת מודלים LLM עם הקשיים רבים יותר. השיטה משתמשת באימון בעל דרגה נמוכה ובקצאות דרגה מותאמות. SlimKV מראה תוצאות טובות בבדיקות LongBench ו-Needle-in-a-Haystack.
תקציר מקורי באנגליתarXiv:2610.02953v1 Announce Type: new Abstract: Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduce per-token KV dimensions but can require full-dimensional reconstruction to apply positional embedding, limiting decoding speedups. We introduce SlimKV, a question-agnostic joint token-feature KV-cache compression method. SlimKV uses low-rank-aware training to compress long contexts into beacon memory states with latent KV representations, together with layer-adaptive rank allocation. We further uncover a positional asymmetry: removing key-side Ro
קרא במקור המקורי