יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ResidualKV: קישור זיכרון KV עם קיצור רזידואלי להקלה באינפרנסה של אורך-תפוקה

ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference
ResidualKV הוא שיטת קיצור זיכרון KV חדשה שמקלה על אינפרנסה של אורך-תפוקה. השיטה מפצלת את הזיכרון KV לקישורים גלובליים וקודים רזידואליים קצוצים. ניתן להקל על הזיכרון והמחשבה על ידי שימוש בשיטה זו. השיטה נבחנה על מודלי LLaMA, QWEN, LLaVA-OV ו-QWEN3-VL.
תקציר מקורי באנגליתarXiv:2602.08005v2 Announce Type: replace-cross Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense of irreversible token eviction, full-cache retention, or full-history reconstruction, limiting their effectiveness for multi-turn interaction and long-form reasoning. Motivated by two empirical properties, Long-Range Inter-Token Similarity and Smooth Residual Distribution, we propose ResidualKV, which factorizes the KV cache into a sparse set of globally retrieved references and compact, quantized residual codes for the remaining tokens. This representation preserves token-specific information without permanent evi
קרא במקור המקורי