יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הערכת מונטה קרלו לגריעת KV Cache

Monte Carlo Estimation for KV Cache Eviction
שיטת LORE-KV משתמשת בהערכת מונטה קרלו כדי לשפר גריעת KV Cache. השיטה מתבססת על דגימת המשך אוטורגרסיבי קצר מהמודל היעד ומשתמשת במצבי השאילתה כדי להעריך את תועלת הטוקנים. השיטה הראתה שיפורים בביצועים על מודלים כמו Qwen2.5-14B ו-Mistral-7B.
תקציר מקורי באנגליתarXiv:2610.07643v1 Announce Type: new Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected
קרא במקור המקורי