כתבה
arXiv cs.LG ·
QUILT: חידוש בהפניית תשומת לב דלילה: רעיון חדש להכנת קוד קודם
QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
QUILT הוא חידוש בהפניית תשומת לב דלילה שמשתמש בביצוע יחד של משאלות סמוכות ושימוש מחדש בנתוני KV. המאמר עוסק באופטימיזציה של הפניית תשומת לב דלילה דרך ביצוע יחד של משאלות סמוכות ושימוש מחדש בנתוני KV.
תקציר מקורי באנגליתarXiv:2610.11134v1 Announce Type: new Abstract: Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its ov
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית