כתבה
arXiv cs.CL ·
השקע ברוחב: סחבת-מדויק לקיצור זמן פענוח בקישורי זיכרון KV בתפיסה ארוכה של חשיבה
Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning
במאמר זה נחקרים סחבת-מדויק בקיצור זמן פענוח בקישורי זיכרון KV בתפיסה ארוכה של חשיבה. המחברים מציגים את BreadthKV, שמשתמש בקיצור זמן פענוח והפליאה של טוקן זיכרון KV. הם מציגים תוצאות של BreadthKV והשוואה לאלגוריתם ThinKV.
תקציר מקורי באנגליתarXiv:2610.05685v2 Announce Type: replace Abstract: Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית