כתבה
arXiv cs.CL ·
האחרון אבל לא האחרון: קליברציה של תשומת לב גבולית להקלה של קאש KV מרוב-תצורה
Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression
במאמר זה, BACON, מתואר כאלגוריתם שמקלה על הקאש KV של מודלי LLM מרוב-תצורה. האלגוריתם נבחן במספר בנקאיים והוכיח תוצאות טובות.
תקציר מקורי באנגליתarXiv:2606.14782v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evidence and discard answer-relevant tokens under aggressive compression. We identify last query attention as a complementary signal for recovering such evidence, though its irrelevant signals may introduce additional noise. We propose BACON, a plug-and-play method that calibrates observation window attention with last query evidence while suppressing noise through intra-layer coherence and inter-layer persistence. Across d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית