כתבה
arXiv cs.LG ·
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
תקציר מקורי באנגליתarXiv:2607.24260v1 Announce Type: new Abstract: Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and confidence signals. Yet LLM serving remains fundamentally oblivious to this rich structure: once such signals are serialized into a prompt, the backend observes only a flat token sequence, forcing dense and uniform consumption of the full key-value (KV) state during decoding. We term this architectural mismatch the Knowledge Selection-Runtime Consumption (KSRC) gap: richer contexts enlarge the full-prompt KV footprint and decode-time memory traffic, increasing latency and degrading throughput even when reasoning depends on only a small fraction of the context. To br
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית