יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

מה עושה דגמים פתוחים נפוצים בייצור — Sujee Maniyam, Nebius

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius
▶ צפה כאן — בלי לצאת מהאתר
דגמים פתוחים כבר דומים לאלה הפרטיים, אבל לשרת אותם במהירות ובעלות נמוכה הוא החלק הקשה. Sujee Maniyam ו-Dylan Bristot מ-Nebius חושפים את האופציות לשרת LLMs פתוחים בייצור.
תקציר מקורי באנגליתOpen models have nearly caught up with proprietary ones. Serving them fast and cheaply is the hard part. Dylan Bristot, who leads product marketing for Nebius Token Factory, and developer advocate Sujee Maniyam explain what it takes to run open LLMs in production. Dylan covers the trade-off between closed APIs and self-hosting, and the inference → data → post-training → deployment loop. Sujee then goes layer by layer through the optimizations behind fast inference: NVFP4 on the latest NVIDIA hardware, engine selection, cache-aware routing, speculative decoding with custom draft models, KV cache offloading, disaggregated prefill and decode, and finding the quantization sweet spot. In this talk: • Why open models are now competitive, and what that means for cost and lock-in • Cache-aware rou
קרא במקור המקורי