כתבה
arXiv cs.LG ·
HeadGuard: הגנה נבחרת לראש של דגימות נמוכ-ביט VLM KV-Cache Quantization
HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
HeadGuard היא שיטת הגנה נבחרת שמגנה ראשי דגימות VLM KV-Cache נמוכ-ביט. היא משפרת את דיוק הדגימות. HeadGuard נבחרה ופותחה על ידי DeepSeek.
תקציר מקורי באנגליתarXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית