יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מנגנוני דגמי זמן-ארוך היברידיים: חלק 1.1: משילוב תשומת העין המלאה למשילוב תשומת העין המשולבת

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
חוקרים חדשים על דגמי זמן-ארוך היברידיים, שמשלבים שונות של תשומת העין, כדי לשפר עמידות לזמן-ארוך וביצועים. המאמר עוסק במנגנוני דגמי זמן-ארוך היברידיים, ובפרט במשילוב תשומת העין המלאה עם תשומת העין המשולבת.
תקציר מקורי באנגליתarXiv:2610.10114v1 Announce Type: new Abstract: The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation.
קרא במקור המקורי