כתבה
arXiv cs.AI ·
AttSVD: דחיסת קאש KV באמצעות SVD-מונחה-תנועה
AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD
AttSVD הוא דחיסת קאש חדשה למודלי תרגומים אוטומטיים, המשמרת כל הטקסט ומחסין אותו באמצעות SVD-מונחה-תנועה.
תקציר מקורי באנגליתarXiv:2610.06927v1 Announce Type: cross Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "interpretable" low-rank compression whose basis is derived from each prompt's own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית