כתבה
arXiv cs.AI ·
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
תקציר מקורי באנגליתarXiv:2610.04753v2 Announce Type: replace-cross Abstract: Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית