כתבה
arXiv cs.CL ·
VestigeKV: הקצאת תשומת לב דלילה במגירה KV
VestigeKV: The NoPE-MLA KV Cache Carries Its Own Sparse-Attention Signal in a Vestigial Branch
VestigeKV: מגירה KV שמפתחת תשומת לב דלילה מתוך התאריך שהיא כבר מחזיקה. המגירה נשארת על ה- GPU ומציעה יתרון במהירות, ולא בזיכרון.
תקציר מקורי באנגליתarXiv:2609.03949v2 Announce Type: replace-cross Abstract: A long-lived KV cache must be compressed before the queries that will read it exist. Selection by observed attention collapses there: on a NoPE-MLA model, H2O and SnapKV retrieve 0.00 and 0.33 of needles at 8x compression, because a token's importance has not yet been observed. VestigeKV instead derives a sparse attention pattern from a signal the cache already carries, occupying the sparse-attention literature's one unoccupied quadrant: training-free and query-independent. In NoPE-MLA the 64-dimensional decoupled branch is a vestige of RoPE that training repurposes into a salience channel; reading 11% of each row, it partitions the cache into an attended tier and a GPU-resident archive that no row ever leaves, reachable each step b
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית