כתבה
arXiv cs.LG ·
Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks
תקציר מקורי באנגליתarXiv:2610.10666v1 Announce Type: new Abstract: Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית