יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תשומת לב דלילה ילידית

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
NAMOH היא מנגנון תשומת לב דלילה ילידית המאפשרת הקטנת נפח הפרמטרים והקשר. היא משתמשת בראשים מרובים המבצעים תשומת לב סיבתית בתוך תת-רצף. NAMOH יכולה לשפר את איכות המודלים ולאפשר הקטנה יעילה של הקשר.
תקציר מקורי באנגליתarXiv:2609.38832v1 Announce Type: new Abstract: Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increa
קרא במקור המקורי