כתבה
arXiv cs.LG ·
RunningTensor: הכללה של תשומת לב ליניארית למצבים רקורסיביים מסדר גבוה
RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
RunningTensor הוא חידוש המרחיב את תשומת הלב הליניארית למצבים רקורסיביים מסדר גבוה. הוא מאפשר ייצוג של אינטראקציות מורכבות יותר במצב, ומשפר את קיבולת הזיכרון העבודה. החידוש הראה תוצאות טובות יותר מאשר מודלים אחרים במטלות שונות.
תקציר מקורי באנגליתarXiv:2609.12814v1 Announce Type: new Abstract: Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order-$o$ tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries. Order $2$ recovers linear attention; we study order $3$ as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length $T$ and improving working memory capacity from $\mathcal{O}(W^2)$ to $\mathcal{O}(W^o)$. On synthetic multi-query associative recall, RunningTensor outperforms linear-attention and SSM baselin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית