כתבה
arXiv cs.LG ·
התקפות חברות-אגן: פגיעה במודלים AI
Membership Inference Attacks for Unseen Classes
אפשרות חדשה להתקפות חברות-אגן: פגיעה במודלים AI שאינם ידועים. חוקרים חשפו חולשה במודלי MIAs הנפוצים, והציעו פתרון חדש.
תקציר מקורי באנגליתarXiv:2506.06488v3 Announce Type: replace Abstract: A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a black-box model. Unfortunately, most \emph{membership inference attacks} (MIAs) used to perform this type of auditing themselves assume \emph{access} to examples of harmful content from the same distribution as the query data. In real-world auditing scenarios, auditors often face legal and ethical restrictions preventing them from accessing a representative set of samples of harmful content to train MIA models effectively. We abstract and formalize this setting into a new data access model, the ``unseen class'' setting, and show that the state of the art MIAs fail
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית