יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

אימות ביצועים של LLM מבוזרים עם NVIDIA srt-slurm

Validating Distributed LLM Serving Benchmarks with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis
NVIDIA מציגה את srt-slurm, כלי לאימות ביצועים של LLM מבוזרים. הכלי מאפשר ליצור תהליכים רבי קנה מידה ולנתח את התוצאות. המאמר מראה כיצד להשתמש בכלי זה עם DeepSeek-R1.
תקציר מקורי באנגליתIn this tutorial, we explore NVIDIA’s srt-slurm framework and learn how we use srtctl to convert declarative YAML configurations into reproducible SLURM benchmark workflows for distributed LLM serving. We set up the project in Google Colab, inspect its internal architecture, define a cluster configuration, dry-run built-in and custom recipes, and model a disaggregated prefill-and-decode deployment for DeepSeek-R1. We also generate parameter sweeps, interact with the typed Python API, validate expanded configurations, and analyze simulated benchmark results through a throughput-versus-latency Pareto frontier. Although Colab does not provide a real SLURM environment, we use it as a practical development workspace to understand, validate, and prepare production-grade benchmark recipes before
קרא במקור המקורי