יום שלישי, 6 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

What Is an Inference Engine, Anyway? — Charles Frye, Modal

▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתA traffic spike on a museum placard generator produces longer waits for the first token and slower tokens afterward. Charles Frye reads those symptoms from an inference dashboard and shows how additional replicas relieve the queue. The example connects an application people can see to the machinery behind its API. He follows a request through server IO, tokenization, scheduling, model execution, and detokenization, using SGLang and vLLM as reference points. The scheduler controls what reaches the GPU and can become a bottleneck even though it performs far less computation. Chatbots, background agents, and document processors place different demands on that machinery, with latency budgets, input and output lengths, and prefix reuse shaping the deployment. Frye then opens up the performance
קרא במקור המקורי