יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MoEless: שירות LLM יעיל עם מומחים ללא MoE

MoEless: Efficient MoE LLM Serving with Serverless Experts
MoEless הוא שירות LLM יעיל שמשתמש במומחים ללא MoE. הוא מורכב על Megatron-LM ומציע זמן עריכה נמוך ועלויות נמוכות יותר.
תקציר מקורי באנגליתarXiv:2603.06350v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proact
קרא במקור המקורי