יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בחינה של שימוש ב- GPU להפנמה של LLM על Nvidia Hopper

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
במאמר זה, נחקר שימוש ב-GPU להפנמה של LLM על Nvidia Hopper. המחברים פרופילו את vLLM עם FlashAttention-3 ו-cuBLASLt על H100 NVL, ומציעים שמונה נקודות תצפית שונות לבחינת שימוש ב-GPU.
תקציר מקורי באנגליתarXiv:2609.12923v1 Announce Type: cross Abstract: A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number. This is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multiplications. On Hopper, the bfloat16 GMMA path executes these operations in fixed 64-row matrix fragments, so small-batch decode can fill only a small fraction of each fragment with real token rows. In this paper, we profile vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length an
קרא במקור המקורי