יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

VFold: חידוש של קיצור זיכרון עבור קודים עם סימטריה

VFold: Symmetry-Aware Cross-Layer Value Cache Compression
המחברים הציגו חידוש לקיצור זיכרון עבור קודים עם סימטריה, כדי לפצל את הזיכרון ולהגדיל את הקונטקסט. החידוש נקרא VFold.
תקציר מקורי באנגליתarXiv:2610.12338v1 Announce Type: cross Abstract: While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method re
קרא במקור המקורי