כתבה
arXiv cs.LG ·
Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data
תקציר מקורי באנגליתarXiv:2606.16952v2 Announce Type: replace Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. However, generating high-utility synthetic data often carries the risk of memorizing and regurgitating private information from the training corpus. In this work, we present a customizable empirical auditing framework designed to detect and explain such data disclosures. Our framework introduces a mechanism to distinguish between "true disclosures"-where the system directly reproduces a user's information-and "phantom disclosures''-where the system incidentally generates a user's data. By partitioning input data into training and holdout sets and applying rigorous
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית