וידאו
YT AI Engineer ·
GPU מת
GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe
▶ צפה כאן — בלי לצאת מהאתר
אימון עצמאי בקנה מידה גדול: Crusoe מציגה את Managed Slurm על Kubernetes. הם מסבירים כיצד AutoClusters טופל לכשל GPU XID 79, ומדגימים כיצד להרוג GPU באמצע אימון ולחזור לשם בתחתית 15 דקות.
תקציר מקורי באנגליתAcross thousands of GPUs, hardware failures are guaranteed. Fixing them by hand at 3 a.m. doesn't scale. Connor Guerrero, Young Jeong and Nikhil Gupta from Crusoe explain why they built Managed Slurm on top of Kubernetes. Slurm gives researchers gang scheduling, topology awareness and familiar sbatch workflows, but falls short on dynamic resources, node health and observability. Kubernetes fills those gaps without either team changing how it works. They walk through AutoClusters' fully automatic remediation of an XID 79 GPU failure, then demo killing a GPU mid-training and resuming from checkpoint in under 15 minutes with no human action. In this talk: • Where traditional Slurm excels for training and where it falls short • Why one stack beats separate Slurm and Kubernetes infrastructure •
קרא במקור המקורי
youtube.com
פתח כתבה מקורית