כתבה
arXiv cs.AI ·
המרכיב הפעיל ב-Muon's Grokking
The Active Ingredient in Muon's Grokking
אופטימייזר Muon מגיע לשיא grokking מהר יותר מ-AdamW. המחקר מציג תוצאות של סריקה רב-זרעים ו- multi-learning-rate, ומצא שהמנגנון העיקרי הוא ה-orthogonalization.
תקציר מקורי באנגליתarXiv:2607.20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW. Prior work attributes this to "spectral-norm constraints plus orthogonalized momentum" but does not isolate which mechanism matters. To better understand Moun's behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect. First, an ablation shows the speedup comes from orthogonalization (the Newton-Schulz iteration): orthogonalize-only matches full Muon, whereas spectral-only is no faster than AdamW and is unreliable, and this verdict holds across learning rates. Second, a mechanistic analysis finds that orthogonalizing optimizers reach generalization at roughly 3x lower spectral norm and, controlling for how muc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית