כתבה
arXiv cs.LG ·
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
תקציר מקורי באנגליתarXiv:2607.24787v2 Announce Type: replace-cross Abstract: Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited accelerator memory. Although expert offloading alleviates memory pressure by moving inactive experts to host memory or storage, it introduces a routing-dependent transfer bottleneck: required experts are known only after native top-\(K\) routing, which serializes routing, expert loading, and expert execution during inference. To address this bottleneck, we propose SpecPrefetch, a parameter-efficient prefetching framework for offloaded MoE inference. SpecPrefetch uses a shared lightweight adapter to predict next-layer expert candidates only for asynchronous
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית