כתבה
arXiv cs.LG ·
Block Sparse Attention with Log-Linear Complexity
תקציר מקורי באנגליתarXiv:2609.31093v1 Announce Type: new Abstract: Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys. Specifically, we construct a coarse-to-fine hierarchy of keys and perform selection from the coarsest level. At each level, LogSumExp scoring is applied to a boun
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית