כתבה
arXiv cs.LG ·
A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
תקציר מקורי באנגליתarXiv:2607.22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examples can be generated at scale but left three questions untested: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet. (i) An independent recognizer proxy drops to ~49% on CIFAR-10 while the model stays at 100% -- a gap a small human pilot (N=5) corroborates directly an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית