כתבה
arXiv cs.LG ·
Page-EntroKV: מנגנון חדש לניהול זיכרון מטמון
Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention
Page-EntroKV הוא מנגנון חדש לניהול זיכרון מטמון במודלים של שפה ארוכי קשב. הוא משתמש במשקלים המבוססים על אנטרופיה כדי לקבוע אילו נתונים לשמור בזיכרון. המנגנון נבדק על מודל Qwen2.5-1.5B-Instruct והראה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2610.03135v1 Announce Type: new Abstract: Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isol
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית