יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

OpWeave: פלטפורמה ניידת להפרדת פעולות למודלי LLM

OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving
OpWeave היא פלטפורמה שמפרידה את שירותי LLM לשלבים קטנים יותר עבור מכשירים הטרוגניים. היא מספקת דגימה ניטרלית של הצבעים של פיצול פעולות ומאפשרת עדכון עצמאי. OpWeave יוצרת תוכניות שמפרידות את הפעולות לשלבים קטנים יותר, ומאפשרת עדכון עצמאי של הפעולות. OpWeave יכולה לצבור עד 1.89 פעמים יותר מאשר הבסיס הטוב ביותר.
תקציר מקורי באנגליתarXiv:2609.14237v1 Announce Type: cross Abstract: LLM serving systems increasingly disaggregate inference into finer-grained stages, with recent approaches separating attention from FFN or MoE execution during decode. This operator-level disaggregated serving (ODS) can improve hardware matching and enable independent scaling, particularly across heterogeneous devices. However, existing systems fix operator boundaries and lack a unified characterization of when disaggregation reduces serving cost. We present OpWeave, an end-to-end framework for heterogeneous ODS. OpWeave provides an analytical cost model that bounds the gains of homogeneous and heterogeneous ODS over colocated serving. It jointly optimizes operator partitioning and deployment configuration through a regularity-aware planner
קרא במקור המקורי