כתבה
arXiv cs.LG ·
שינוי תנועה לינארית
Switching Linear Attention
אחד האתגרים המרכזיים בלמידת מכונה הוא עיצוב שכבות תזמון רציפות עם ניצול יעיל של זיכרון. SwiLA, שינוי תנועה לינארית, מציעה שכבה חדשה שמשלבת יכולת תיאור גבוהה עם זיכרון קבוע. SwiLA מציגה יכולת תיאור גבוהה יותר וניצול יעיל של זיכרון, ומציעה פתרון חדש לאתגר זה. SwiLA מבוססת על פרוטוקול עקיפה זמני, שמאפשרת עדכון זמני של המצב. SwiLA מציעה פתרון חדש לאתגר זה, ומציעה יכולת תיאור גבוהה יותר וניצול יעיל של זיכרון.
תקציר מקורי באנגליתarXiv:2609.39034v1 Announce Type: new Abstract: Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-tim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית