כתבה
arXiv cs.AI ·
הקפאת המפרש, תיקון המפרש: תיקון-פרמטרים-קצר-מועיל להתאמה להקטנת זיכרון KV-קאש
Freeze the Decoder, Heal the Encoder: Parameter-Efficient Adaptation for SVD-Based KV-Cache Compression
אפליקציה חדשה להתאמה קצר-מועיל להקטנת זיכרון KV-קאש, המשתמשת ב-SVD ופינוי פרמטרים. התאמה זו נותנת תוצאות טובות יותר וחסכוניות בזיכרון.
תקציר מקורי באנגליתarXiv:2610.10552v1 Announce Type: cross Abstract: Comparing parameter-efficient fine-tuning recipes under a single, shared learning rate is a common but flawed practice: when the arms being compared have very different trainable-parameter counts, a shared rate can simultaneously depress the larger arms' means and inflate their variance, manufacturing a large, seemingly multi-seed-significant advantage for the smallest arm that is not a real effect. We document this confound in a concrete setting: post-hoc SVD-based KV-cache compression, where an already-pretrained model is converted to a low-rank (multi-head-latent-attention-style) cache by factorizing its key/value weights into a down-projection ("encoder") and an up-projection ("decoder"), after which a short fine-tune ("healing") recove
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית