כתבה
arXiv cs.LG ·
EAServe: שירות ניתוח-מודולרי ל-LLMs
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
EAServe משפר את שירות ה-LLM המודולרי עם ניתוח-מודולרי והפרדת GPU דינמית. הוא מציע פתרון לבעיית השימוש הלא-אופטימלי של GPU בשלב ה-Encode של LLMs.
תקציר מקורי באנגליתarXiv:2609.31551v1 Announce Type: cross Abstract: Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית