יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ResidualQuant: קוונטיזציה ל-KV Cache

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
ResidualQuant הוא שיטה חדשה לקוונטיזציה של KV Cache, המשפרת את יעילות הזיכרון והקצב במודלים של Looped Transformers. השיטה משתמשת במצבים הסופיים של KV כרפרנס ומייצגת את הלולאות הנותרות באמצעות שאריות בנקודה נמוכה. ResidualQuant מאפשרת קוונטיזציה מדויקת עד INT2 תוך כדי שמירה על יעילות גבוהה.
תקציר מקורי באנגליתarXiv:2610.10381v1 Announce Type: new Abstract: Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops wit
קרא במקור המקורי