כתבה
arXiv cs.AI ·
Monte Carlo Estimation for KV Cache Eviction
תקציר מקורי באנגליתarXiv:2610.07643v1 Announce Type: cross Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projecte
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית