יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

FluidPD: גמישות במקום לשרת LLM מפוצל

FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
FluidPD הוא מערכת שרת LLM מפוצלת המספקת גמישות במקום. היא משפרת את עמידת ה-SLO על ידי העברת חישובים לעובדים פנויים. המערכת משתמשת במדדי לחץ קלים כדי לזהות לחץ משאבים לפני הפרות SLO.
תקציר מקורי באנגליתarXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated
קרא במקור המקורי