כתבה
arXiv cs.CL ·
MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
תקציר מקורי באנגליתarXiv:2609.31261v1 Announce Type: new Abstract: The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where positional relevance can decay and where broader interactions must be preserved. We introduce Mixture of Semantic Attention Regimes (MoSAR), which learns such an adaptive, controlled-decay geometry over query--key interactions. Input-conditioned query and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית