כתבה
arXiv cs.LG ·
Hidden Gauge Controls Feature Specialization in ReLU Networks
תקציר מקורי באנגליתarXiv:2608.06766v2 Announce Type: replace Abstract: The success of deep learning depends on learning useful representations, yet predicting how training organizes these representations across neurons remains difficult. In this work, we show that changing the scale of initial weights can determine which neurons learn a feature without altering any neuron's initial contribution. We construct ReLU networks with identical initial features and predictions that reach the same final predictions with different roles for their neurons. In one, all neurons share the learned feature. In the other, one neuron acquires it while every other neuron's contribution vanishes. The only change is the relative scale of each neuron's input and output weights. Our analysis explains how an initial learning advant
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית