כתבה
arXiv cs.AI ·
הקביעות פורצות את המהירות הלמידה
Equivariance Breaks the Learning Rate
הקביעות פורצות את המהירות הלמידה. חוקרים גילו כי רשתות קביעות נלמדות בקצב גבוה יותר עם אופטימיזציה של Muon, וזאת כתוצאה משינוי בקביעות הגרדיאנט.
תקציר מקורי באנגליתarXiv:2609.08381v2 Announce Type: replace-cross Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix structured optimizers such as Muon can perform better, with the reasons for these gains only partly understood. We identify one source of this difference inside equivariant layers. An equivariant layer learns one channel mixing matrix $W_l$ per degree $l$, which we call an irrep block, and shares it across the $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. This sharing sums gradient contributions across components and can produce different update scales under SGD. Adam's entrywise normalization reduces sensitivity to gradient scale, but neither optimizer directly controls the effective step size of each block. A single learni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית