כתבה
arXiv cs.AI ·
KV-Kaizen: למידת קביעת צמצום קאיזן להצפיית קודקס
KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
KV-Kaizen: פיתוח קביעת צמצום קאיזן להצפיית קודקס, המאפשרת צמצום זיכרון ב-LLMs. השיטה משתמשת בלמידת מכונה כדי לבחור את הקביעות הטובות ביותר לצמצום קודקס, תוך שמירה על דיוק. KV-Kaizen יכולה להצטרף לשיטות אחרות לצמצום קודקס, כגון פינוי, ולהגדיל את היעילות שלהן.
תקציר מקורי באנגליתarXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית