כתבה
arXiv cs.LG ·
Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
תקציר מקורי באנגליתarXiv:2610.00136v1 Announce Type: cross Abstract: This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilon$ = 0.15 and 1.72% at $\epsilon$ = 0.30; under PGD it fell to 32.47% and 0.41%, and a bit-depth-reduction defense recovered only part of the loss. In the second experiment, DistilBERT fine-tuned on the SMS Spam Collection reached 98.75% accuracy and a 94.96% F1-score, but a controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed ex
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית