כתבה
arXiv cs.CL ·
CoSA: זירוז אינפרנס לאורך הקשב דרך תשומת לב דלילה משותפת
CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
CoSA היא שיטה חדשה לזירוז אינפרנס לאורך הקשב על ידי שימוש בתשומת לב דלילה משותפת. היא משלבת פרוקסי תבוני עם ליבת סדר מוגדר. CoSA מאפשרת להגיע לרמות דיוק גבוהות יותר עם תקציבים נמוכים יותר.
תקציר מקורי באנגליתarXiv:2607.25291v1 Announce Type: new Abstract: The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blocks, while the kernel can only apply the sparse mask mechanically, leading to an evident drop in model accuracy. We propose CoSA, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK). In the first stag
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית