יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרישת שקילות: כיצד הקצאת קצב הלמידה נשברת

Equivariance Breaks the Learning Rate
פרישת שקילות: כיצד הקצאת קצב הלמידה נשברת. חוקרים גילו ש-Adam נכשל באופן שאינו ניתן להסבר, ופיתחו פתרון לבעיה זו.
תקציר מקורי באנגליתarXiv:2609.08381v1 Announce Type: new Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. For a single application of the layer, the gradient of $W_l$ sums $2l+1$ outer product contributions and has rank at most $2l+1$. Adam rescales stored weights individually without using the irrep boundaries, so one learning rate can produce different spectral step sizes across blocks within a layer. We address this mismatch by normalizing each b
קרא במקור המקורי