יום שבת, 10 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

מ-15% ל-90% GPU: תיקון פייפליין, ולא את המודל

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model
▶ צפה כאן — בלי לצאת מהאתר
פתרון בעיות בפייפליין הנתונים יכול לשפר את השימוש ב-GPU באופן ניכר.
תקציר מקורי באנגליתYour GPUs might be waiting on your data, not your model. This pipeline went from 15% to 90% GPU utilization. Tarun Sunkaraneni from Amazon AGI shows that the hardest part of fast multimodal training often isn't the GPU at all, but keeping it fed with data. Training a Qwen3-VL-style model on images stored in S3, the baseline pipeline spent about 85% of its time waiting on data, mostly loading, decoding and resizing images one at a time. He fixes the bottlenecks one by one: concurrency with asyncio and Ray actors, prefetching so data is ready before the trainer asks, and Ray's object store to stop copying big image tensors between processes. Then, at production scale, a new bottleneck appears on a single machine's network card, solved by spreading workers across nodes and using zero-copy rea
קרא במקור המקורי