כתבה
arXiv cs.AI ·
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
תקציר מקורי באנגליתarXiv:2607.24887v1 Announce Type: cross Abstract: Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remains beneficial on unseen data. We study this problem for function-preserving residual expansion and introduce the effective alignment dimension, a measurable quantity describing the signal-noise geometry of activation gradients. By deriving the exact mean and variance of the inner product between independently estimated training and test gradients, we obtain a finite-sample upper bound on misalignment probability. The bound depends only on the effective alignment dimension and an effective sample size, requiring finite second moments and a nonzero population gradien
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית