כתבה
arXiv cs.LG ·
Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
תקציר מקורי באנגליתarXiv:2609.11744v1 Announce Type: cross Abstract: Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית