כתבה
arXiv cs.CL ·
תשומת לב דלילה מהירה
Block Sparse Flash Attention
Block Sparse Flash Attention היא שיטה חדשה לחישוב תשומת לב דלילה, המאיצה את תהליך ההסקה במודלים גדולים. השיטה נבדקה על מודל Llama-3.1-8B והראתה שיפור בביצועים.
תקציר מקורי באנגליתarXiv:2512.07011v2 Announce Type: replace-cross Abstract: Modern large language models increasingly require long contexts for reasoning and multi-document tasks, but attention's quadratic complexity creates a severe computational bottleneck. We present Block Sparse Flash Attention (BSFA), a drop-in replacement that accelerates long-context inference while preserving model quality. Unlike methods that predict importance before computing scores, BSFA computes exact query-key similarities to select the top-k most important value blocks for each query. By comparing per-block maximum scores against calibrated thresholds, we skip approximately 50% of the computation and memory transfers for pruned blocks. Our training-free approach requires only a one-time threshold calibration on a small datase
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית