יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration

תקציר מקורי באנגליתarXiv:2610.11214v1 Announce Type: cross Abstract: KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Und
קרא במקור המקורי