יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תיקוני תדירות קטנים יכולים לשנות את הדברים שנשמרים בהתקפת קיצור

Small Frequency Corrections Can Change What Survives KV Cache Compression
תיקוני תדירות קטנים יכולים לשנות את הדברים שנשמרים בהתקפת קיצור. זה נכון גם למודל Llama-3.2-1B.
תקציר מקורי באנגליתarXiv:2608.27128v4 Announce Type: replace Abstract: Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value energy measures entry strength but does not distinguish isolated keys from those with many similar neighbors. We introduce TwinKV, a training-free method that discounts value energy by nonlocal post-RoPE key frequency. Prefix attention allocates head capacities, while retained entries preserve their original keys and values under an exact storage budget. Across four language models, TwinKV exceeds five evaluated compressed baselines in mean score on LongBench, LooGLE, and RULER at 50\% KV removal. Component controls isolate the frequency contribution. On Llama-3.2-1B RULER at 75\% removal, norm
קרא במקור המקורי