יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימון מורכב של מומחים: פיתוח רובוסטנטי

Distributionally Robust Mixture-of-Experts Training
אימון מורכב של מומחים: פיתוח רובוסטנטי. DRMoET, פיתוח רובוסטנטי של MoE, משפר את הביצועים של MoE בעזרת אובייקטיב חדש.
תקציר מקורי באנגליתarXiv:2610.07207v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-M
קרא במקור המקורי