יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

תקציר מקורי באנגליתarXiv:2606.21848v2 Announce Type: replace-cross Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability limitations of the standard QKV attention mechanism. The Key-Value (KV) cache is a major bottleneck during long-context inference. We propose Keyless Attention, a novel attention mechanism that introduces a dedicated value-space routing projection to replace the conventional key projection, thereby eliminating key representations from the attention computation. This design yields a Value-Only Cache that reduces KV-cache memory and access overhead by 50% compared with standard attention while improving decode throughput. Experiments across five models and four architectures show that Keyless
קרא במקור המקורי