כתבה
arXiv cs.LG ·
BeaconKV: קישור-ערך: דחיסה מונחה על-ידי שאילתות Beacon למערכת זיכרון קווים-ערכים לאינפרנסה יעילה של דגמי סיבוכיות גדולים
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
מאמר חדש מציג שיטה חדשה לאינפרנסה יעילה של דגמי סיבוכיות גדולים, על ידי דחיסת קישור-ערך. השיטה, בשם BeaconKV, משתמשת בשאילתות Beacon כדי לקבוע מה יהיה צורך להזיכרון. המאמר מציג תוצאות טובות יותר משיטות דחיסה אחרות.
תקציר מקורי באנגליתarXiv:2609.04971v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית