כתבה
arXiv cs.LG ·
SMELT: חוקי גדילה למודלי MoE Looped Transformers
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
חוקי גדילה למודלי MoE Looped Transformers. המחברים חקרו את השימוש במודלי MoE Looped Transformers, ופיתחו רצפה חדשה שנקראת SMELT. SMELT משתמשת בטכניקה של looping, שמאפשרת למודל לחזור על עצמו ולשפר את התוצאות. המחברים גילו ש-SMELT משפרת את התוצאות ב-6.8-18.0% בהשוואה למודלים קודמים.
תקציר מקורי באנגליתarXiv:2609.01343v2 Announce Type: replace Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית