כתבה
arXiv cs.CL ·
MoE$^2$-LoRA: היכרות בין מודלים MoE לעיבוד LoRA
MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
MoE$^2$-LoRA הוא שיטה חדשה לעיבוד דגמי MoE. השיטה משלבת את היתרונות של עיבוד LoRA עם דגמי MoE, תוך שימוש במודול RCP כדי לשפר את יעילות העיבוד. השיטה הוכחה כיעילה במגוון ניסויים.
תקציר מקורי באנגליתarXiv:2607.21978v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models remains underexplored. Existing PEFT methods for MoE either ignore router priors with uniform adapters, reducing efficiency and risking forgetting, or rely on static expert selection, limiting per-token capacity and cross-expert feature learning. In this paper, we make the first attempt to fine-tune MoE models with MoE-style low-rank adaptation: our method, entitled MoE$^2$-LoRA, deeply couples the pretrained expert specialization with task-specific adaptivity via a dual-channel Routing-Conditioned Projection (RCP) module, which reuses base router activations to inform LoRA routing. We further int
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית