כתבה
arXiv cs.LG ·
בחירת חוקרים בצורה זיהוים-מודעת-צרכן לצורך דקודינג מורכב של MoE
BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding
בחירת חוקרים בצורה זיהוים-מודעת-צרכן לצורך דקודינג מורכב של MoE. ניתן לשפר את האיזון בין דיוק לביצועים בלי לאמן מחדש.
תקציר מקורי באנגליתarXiv:2609.36222v1 Announce Type: new Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert weights to be transferred. Most expert-reduction techniques make retention decisions independently for each token and therefore do not address this batc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית