כתבה
arXiv cs.LG ·
MaskCoFT: שיפור יעילות זיכרון במודלים MoE
MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
MaskCoFT הוא שיטה חדשה לשיפור יעילות זיכרון במודלים MoE. היא מאפשרת למודלים להשתמש בפחות זיכרון ולהגיע לתוצאות טובות יותר. השיטה נוסתה על מודלים כמו DeepSeek והראתה שיפורים משמעותיים.
תקציר מקורי באנגליתarXiv:2609.34077v2 Announce Type: replace Abstract: Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the token
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית