כתבה
arXiv cs.AI ·
הגברת תפוקת ה-CPU לשיחה: אופטימיזציה חד-פעמית לביצועי יישור
Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing
אופטימיזציה של תפוקת CPU לשיחה על פלטפורמות שרות-ללא-שרת. המחברים פיתחו שיטה לאופטימיזציה של CPU-seconds ו-GB-seconds, והציגו תוצאות של 2.71 אלפי ננו-שניות ל-CPU-שנייה, ו-0.0153 דולר לשעה, כ-4.1 פעמים פחות מהתוצאות של PyTorch.
תקציר מקורי באנגליתarXiv:2610.00063v1 Announce Type: cross Abstract: Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or latency alone. Conventional runtimes are poorly suited to this setting: per-request parallelism causes CPU contention under concurrency, while warm instances retain gigabytes of billable inference and page-cache state. We address these costs with request-sized concurrent inference, which bounds per-request CPU parallelism, and a reclaimable instance lifecycle, which releases inference state and page-cache memory after idle periods while retaining th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית