כתבה
arXiv cs.AI ·
שימור סטרוקטורה של רוטר בפוסט-אימפלמנטציה של MoE
Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training
במאמר זה, המחברים חוקרים את השפעת שימור סטרוקטורה של רוטר בפוסט-אימפלמנטציה של MoE על יכולת המודל להתאים לתחום. הם מציגים טכניקה חדשה של Router Prior Bias (RPB) שמאפשרת שימור זה בצורה רכה. התוצאות המוצגות במאמר חושפות את יעילות RPB בשיפור יכולת המודל להתאים לתחום.
תקציר מקורי באנגליתarXiv:2609.08115v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training inherits a different situation: the base router already encodes non-uniform expert co-activation structure, which a re-imposed uniformity objective flattens away. We show that downstream performance depends instead on holding this inherited routing softly, a principle we term soft router anchoring, and instantiate it as Router Prior Bias (RPB), a training-time bias that pulls the router logits toward a prior read off the frozen base router while leaving the router itself trainable. On math post-training of Moonlight-16B-A3B, RPB attains 45.77 in-domain accuracy against 31.91 under re-applied LB
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית