כתבה
arXiv cs.AI ·
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
תקציר מקורי באנגליתarXiv:2609.38166v1 Announce Type: cross Abstract: Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית