כתבה
arXiv cs.LG ·
PAC-Private Autoregressive Generation
PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement
PAC-Private Autoregressive Generation הוא שיטה חדשה להגנה על פרטיות במודלים של שפה. השיטה משתמשת ברעש כדי להגן על המידע הפרטי. המחברים בדקו את השיטה על מודל GPT-2-small וקבלו תוצאות טובות.
תקציר מקורי באנגליתarXiv:2609.05676v1 Announce Type: new Abstract: Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. Private prediction protects these releases. Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons. PAC privacy instead calibrates noise to output variability across possible secrets, adding less noise when predictions are stable. To our knowledge, PAC-private prediction has not previously been extended from classification to autoregressive generation. We construct $m=128$ overlapping worlds from the private corpus, with each record appearing in exactly $m/2$ worlds, and train one adapter per world over a frozen public model. The
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית