כתבה
arXiv cs.CL ·
Read What Matters: Query-Adaptive Quantization for KV Caches
תקציר מקורי באנגליתarXiv:2610.11245v1 Announce Type: cross Abstract: KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית