כתבה
arXiv cs.LG ·
Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
תקציר מקורי באנגליתarXiv:2609.30738v1 Announce Type: cross Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $\lambda_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary with depth, we compare fixed global coefficients with three-segment and quadratic profiles. Only these depth profiles are searched on a development split under a $\sinh$ reparameterization. On all 16 English LongBench datasets with Mistral-7B at a bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית