יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימות שחזור מטמון מצב היברידי ל-GLM-5.3-Flash

Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache
חוקרים בדקו שחזור מטמון מצב היברידי עבור מודל GLM-5.3-Flash. הם השתמשו בנקודת אחז קוונטיזציה RedHatAI/GLM-5.3-Flash-NVFP4 עם vLLM ו-LMCache. התוצאות הראו שיפור ביצועים ואמינות.
תקציר מקורי באנגליתarXiv:2609.15030v1 Announce Type: cross Abstract: External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache under four-way tensor parallelism. A complete-hit recovery mismatch restored state for the full prompt while the scheduler credited one fewer token. We aligned recovery through strict-prefix lookup and established a numerical comparison using shared computation corrections, matched checkpoint scheduling, and fixed per-rank kernel configurations. In a nine-length serial workload, agreement with the modified recomputation control improved from 34/36 to 36/36 generations, each containing 64 token IDs. A separate
קרא במקור המקורי