יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תשומת ליניארית מתחלפת

Switching Linear Attention
תשומת ליניארית מתחלפת (SwiLA) היא שכבת רצף חדשה המאפשרת יעילות וביצועים טובים יותר. SwiLA משלבת בין תשומת ליניארית לבין תשומת softmax, ומראה ביצועים טובים במבחני זיכרון אסוציאטיבי, למידת שפה ודגמי שפה.
תקציר מקורי באנגליתarXiv:2609.39034v1 Announce Type: cross Abstract: Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-t
קרא במקור המקורי