כתבה
arXiv cs.CL ·
שיקום איכות למטמון KV מקוטע
Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation
חוקרים מציגים שיטה לשיקום איכות מטמון KV מקוטע. השיטה משתמשת בעידון תשומת לב בדרגה נמוכה. הניסויים בוצעו על מודלים כמו TinyLlama-1.1B ו-Gemma-4-12B.
תקציר מקורי באנגליתarXiv:2609.04263v1 Announce Type: cross Abstract: Low-bit key--value (KV) caches reduce the memory required for autoregressive decoding, but the resulting quality loss depends on the model and quantizer. We keep the quantizer fixed and distill the floating-cache model's behavior into low-rank Q/K/V projection updates while the student executes a physically packed incremental cache. Across three seeds, 4-bit affine-cache adapters recover $54.24\%\pm2.47\%$ of the held-out perplexity gap on TinyLlama-1.1B and $75.96\%\pm4.04\%$ on Gemma-4-12B. On the same frozen NF4 Llama-3.1-8B base, one validation-selected run per quantizer recovers $60.42\%$ under KIVI K2V2 and $37.61\%$ under KVarN K4V2, while preserving 180-case associative retrieval. Gemma's score on an official 4K/8K RULER subset rise
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית