כתבה
arXiv cs.AI ·
LOCKS: סיכומי ספקטרליים לקיצור סיכומי חותם-מפתח לדקודים באורך-קשר
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
המאמר LOCKS מציג סיכומי ספקטרליים שמקצרים את סיכומי החותם-מפתח לדקודים באורך-קשר. הסיכומים מאפשרים דקודים באורך-קשר עם זיכרון-קווי (KV) קטן יותר. המאמר מציג תוצאות של LOCKS על ידי השוואה לדקודים עם KV גדול יותר.
תקציר מקורי באנגליתarXiv:2607.24555v3 Announce Type: replace-cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-specific directions; fitting a basis to each page better identifies the pages receiving the most attention at comparable stored selector cost. LOCKS stores a rank-$r$ spectral summary per page, reconstructs its within-page logits, and selects pages by log-sum-exp mass without reading candidate keys or values. It stays within about a point of FullKV on LongBench-v1, tracks the read-every-key exact-LSE oracle on RULER down to the smallest budgets, and retains quality furthest unde
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית