יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CacheReforge: שיקום מטמון KV

CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
CacheReforge הוא אלגוריתם שמשפר את השימוש במטמון KV במודלים גדולים. הוא מודד את השינויים במודל ומחליט האם לבצע חישוב מחדש או להשתמש בנתונים הקיימים. האלגוריתם נבדק על מודל Qwen2.5-1.5B ו-Qwen2.5-7B והראה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2609.30884v1 Announce Type: new Abstract: Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from earlier adapter versions nor distinguish update propagation from the recomputation required for behavioral recovery. To address these gaps, we introduce CacheReforge, which represents stale KV caches as layerwise mixed-version objects. It combines per-
קרא במקור המקורי