כתבה
arXiv cs.LG ·
חוקי הקנה של היפרפרמטרים במודלים MoE דלילים
Hyperparameter Scaling Laws Across MoE Sparsity
חוקרים גילו חוקי הקנה חדשים להיפרפרמטרים במודלים MoE דלילים. המחקר בדק 1800 ריצות ומצא שהקצב האופטימלי וגודל הבאטץ' משתנים עם יחס ההפעלה. התוצאות מאפשרות העברה נאותה של היפרפרמטרים בין רמות דלילות שונות.
תקציר מקורי באנגליתarXiv:2609.08690v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית