כתבה
arXiv cs.LG ·
תשומת לב גמישה עם סף
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
חוקרים מציגים אלגוריתם תשומת לב גמיש עם סף, המשפר את יעילות הפענוח במודלים עם הקשר ארוך. האלגוריתם מאפשר קיצוץ זיכרון משמעותי תוך שמירה על איכות הפלט.
תקציר מקורי באנגליתarXiv:2609.20888v2 Announce Type: replace Abstract: Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding. ETA predicts dynamic, contextual thresholds directly from query representations, adjusting context retention depending on the task at hand. To learn this policy from scratch while enabling near lossless KV cache pruning at inference time, we filter attention logits through a shifted SiLU gate during training. We show theoretically and empirically that this creates a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית