יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שליטה בעוצמה וזמן בשירות LLM: פירוק-שלבי והתאמה למודל

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
במאמר זה, צוות מחקר עוסק בשיפור שירות LLM על ידי שליטה בעוצמה וזמן. הם מציג פירוק-שלבי והתאמה למודל, שמאפשרת שיפור של 20.4% בעוצמה ו-3.5% בזמן. המחקר נערך על ידי NVIDIA ומציג תוצאות טובות יותר מאשר פרופילי עוצמה קיימים.
תקציר מקורי באנגליתarXiv:2609.11133v1 Announce Type: new Abstract: Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to prefill and decode GPUs that operate in opposite hardware regimes. We hypothesize that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely re
קרא במקור המקורי