כתבה
arXiv cs.LG ·
קוונציזציה: גורם לפיצול באפליות קצר-זיכרון ב-LLM Serving
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving
קוונציזציה של המודל גורמת לפיצול באפליות קצר-זיכרון ב-LLM Serving. המחקר מצא שקוונציזציה של 16-bit גורמת לפיצול ב-36.2% מהפרקים, ושקוונציזציה של 4-bit גורמת לפיצול ב-75.0%.
תקציר מקורי באנגליתarXiv:2609.04748v1 Announce Type: cross Abstract: Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent optimization. We measure what it costs in reproducibility, and find that the cost rises sharply with weight quantization. Holding the model, decoding parameters, seed, and request order fixed, and issuing every request serially at batch size one, we ran an eighty-episode multi-turn agentic tool-use workload with caching enabled and disabled across two engines and four weight formats. Enabling the cache changed the agent's trajectory on 36.2 percent of episodes at 16-bit precision and on 75.0 percent at four-bit, a gradient that survives re-mea
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית