כתבה
arXiv cs.LG ·
גבולות רדמאכר צפופים לרשתות נוירונים דלילות
Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks
חוקרים פיתחו גבולות רדמאכר צפופים עבור רשתות נוירונים דלילות. המחקר בוחן את הסיבוכיות הסטטיסטית של רשתות עם הפעלה דלילה. התוצאות מראות כי הגבולות החדשים משפרים את ההבנה שלנו על התנהגות רשתות נוירונים.
תקציר מקורי באנגליתarXiv:2609.09130v1 Announce Type: new Abstract: An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית