כתבה
arXiv cs.LG ·
SoK: Privacy Attacks on Machine Learning via Explainable AI
תקציר מקורי באנגליתarXiv:2609.10627v1 Announce Type: cross Abstract: Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership inference, and model inversion, treating attribute inference as partial inversion. Existing work is often labeled only black- or white-box, obscuring substantial differences in what explanation signal reaches an adversary. We therefore separate model knowledge from explanation acquisition and identify five paths: target-released, attacker-derived, secondary disclosure, privileged access, and released global artifacts. Across these paths, explanations reduce extraction cost, expose membership signals through explanation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית