כתבה
arXiv cs.LG ·
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache
תקציר מקורי באנגליתarXiv:2609.15030v1 Announce Type: cross Abstract: External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache under four-way tensor parallelism. A complete-hit recovery mismatch restored state for the full prompt while the scheduler credited one fewer token. We aligned recovery through strict-prefix lookup and established a numerical comparison using shared computation corrections, matched checkpoint scheduling, and fixed per-rank kernel configurations. In a nine-length serial workload, agreement with the modified recomputation control improved from 34/36 to 36/36 generations, each containing 64 token IDs. A separate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית