כתבה
arXiv cs.CL ·
SGD-KV: דחיסת קאש KV תחת הנחיית סיכום
SGD-KV: Summarization Guided KV Cache Compression
מחקר חדש מציע שיטה לדחיסת קאש KV במודלי שפה גדולים, ומציג תוצאות טובות במבחנים שונים.
תקציר מקורי באנגליתarXiv:2609.03235v1 Announce Type: new Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We present SGD-KV (Summarization-Guided KV Cache Compression), a head-aware framework that leverages a novel chunk-summarization diagnostic task to systematically identify and prioritize attention heads specialized in hierarchical information aggregation. Experiments on Qwen2.5-7B-1M and Qwen3-32B across diverse long-context benchmarks demonstrate that SGD-KV achieves state-of-the-art performance with contexts up to 1M tokens, while reducing KV cache me
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית