כתבה
arXiv cs.LG ·
CommunityKV: פיענוח ארוך-הקשר יעיל
CommunityKV: Efficient Long-Context Decoding via Graph Partitioning
CommunityKV הוא כלי לפיענוח ארוך-הקשר יעיל יותר. הוא משתמש בגרף קהילות לאיתור קבוצות טוקנים סמנטיות. נבדק על מודלים Qwen3 ו-Llama-3.1.
תקציר מקורי באנגליתarXiv:2610.00418v1 Announce Type: new Abstract: Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local u
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית