יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אופטימיזציה של AI Inference ברחבי המערכת להפצה

Optimizing AI Inference Across the Deployment Stack
המאמר עוסק באופטימיזציה של AI Inference ברחבי המערכת להפצה, כולל טכניקות שונות ומערכות. המחברים חוקרים את ההשפעה של טכניקות שונות על תפעול המערכת. המאמר כולל ניתוחים ומודלים שונים, כולל רופליין ומודלי תורנות.
תקציר מקורי באנגליתarXiv:2609.10550v1 Announce Type: cross Abstract: AI deployment performance is shaped not by model architecture alone, but by interactions among compression, compiler transformations, and serving policies. Published benchmarks often report latency and throughput under incomparable conditions, limiting their use for deployment decisions. This paper presents a unified analytical treatment of inference optimization across the deployment stack. We introduce a three-layer taxonomy covering model-level techniques such as quantization, pruning, and distillation; compiler transformations such as graph fusion, layout optimization, and kernel autotuning; and system policies such as dynamic batching, admission control, and memory tiering. We formulate deployment as a constrained multi-objective optim
קרא במקור המקורי