יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תשומת לב דלילה היא קירוב מטריצה, לא בחירה מתוך שקית ערכים

Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
חוקרים הציעו שיטה חדשה לתשומת לב דלילה, Matrix Approximation Sparse Attention (MASA), המתבססת על קירוב מטריצה ולא על בחירה של ערכים גדולים. MASA משפרת את דיוק המודלים ויעילותם.
תקציר מקורי באנגליתarXiv:2610.10871v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated a
קרא במקור המקורי