וידאו
YT AI Engineer ·
Is Speculative Decoding Worth It? Profiling vLLM on NVIDIA Blackwell — Akamai
▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתA small model guesses the next few tokens. The big model checks them in one pass. When is that worth it? Sheilah Kirui, developer advocate at Akamai, explains speculative decoding and how to tell whether it's worth turning on. She walks through the prefill and decode phases of inference, how a small draft model proposes tokens that the target model verifies in a single forward pass, and the memory cost of hosting a second model and its KV cache. She covers how to choose a draft model, then demos both setups side by side with vLLM on a single NVIDIA Blackwell GPU: structured output ran 1.6x faster with a high acceptance rate, while creative writing saw far lower acceptance. She also explains why long-context and high-concurrency workloads gain less. In this talk: • Prefill vs decode, and wh
קרא במקור המקורי
youtube.com
פתח כתבה מקורית