כתבה
arXiv cs.AI ·
ValueDiff: קצירת קאש KV על-פי גאומטריה של ערכים ל-LLMs
ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
ValueDiff: קצירת קאש חדשה, המשתמשת בגאומטריה של ערכים, ל-LLMs שמקצירים סינך. השיטה נבחנה בשלושה בסיסים: RULER, LongBench ו-MATH-500. ValueDiff הציגה תוצאות טובות יותר משיטות קצירה קודמות, כולל זו המשתמשת בקאש צפופה.
תקציר מקורי באנגליתarXiv:2609.23314v2 Announce Type: replace-cross Abstract: Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית