כתבה
arXiv cs.LG ·
Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus
תקציר מקורי באנגליתarXiv:2609.39408v1 Announce Type: new Abstract: Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in $H^1(\gamma)$, we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank-$r$ teacher subspace and the leading $r$-dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least $1/2$, and the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית