כתבה
arXiv cs.CL ·
הפעלות מסיביות במודלים גדולים של שפה
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
חוקרים חשפו הפעלות מסיביות במודלים גדולים של שפה עם תשומת לב ליניארית היברידית. המחקר בוחן את המבנה, התפתחות והמשמעות התפקודית של הפעלות הללו. התוצאות מראות כי הפעלות המסיביות מופיעות במבנים שונים ויכולות להשפיע על ביצועי המודל.
תקציר מקורי באנגליתarXiv:2608.12149v3 Announce Type: replace Abstract: We present the first systematic study of massive activations (MAs) in layer-interleaved Hybrid linear attention large language models (HLA LLMs), examining their architectural organization, training-time emergence, underlying mechanisms, and functional significance. Across five linear attention architectures, six hybridization configurations, and five input domains, we identify two architecture-aligned morphologies: pre-attention spikes (PAS) immediately before full attention and inter-spike plateaus (ISP) persisting through intervening linear attention layers. Denser full attention increasingly connects PAS through ISP, approaching the persistent MAs of conventional Transformers. This organization also recurs across 12 public checkpoints
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית