כתבה
arXiv cs.AI ·
הזרקת דברים: כיבוי יכולת בדיבור טקסט-קול
Speech Generation Speaker Poisoning: Capability Erasure in Zero-Shot Text-to-Speech
מערכות TTS חסרות דוגמאות יכולות להדגיש קולות, אך אנחנו יכולים למנוע זיהוי זהויות דוברים מטרה מלהיות מוסוות. המחקר חוקר כיצד למנוע זיהוי זהויות דוברים מטרה במערכות TTS חסרות דוגמאות. התוצאות מציגות כיצד למנוע זיהוי זהויות דוברים מטרה במערכות TTS חסרות דוגמאות. המחקר חוקר כיצד למנוע זיהוי זהויות דוברים מטרה במערכות TTS חסרות דוגמאות.
תקציר מקורי באנגליתarXiv:2603.07551v3 Announce Type: replace-cross Abstract: Recent zero-shot Text-to-Speech (TTS) systems can clone previously unseen voices from only a few seconds of audio. We formulate Speech Generation Speaker Poisoning (SGSP), a task that seeks to prevent a model from synthesizing targeted speaker identities while maintaining performance on all other speakers. Unlike conventional machine unlearning, removing training examples is insufficient because modern zero-shot TTS systems can reconstruct identities through learned speaker representations and strong generalization capabilities. We evaluate both inference-time filtering and parameter-modification approaches across settings involving 1, 15, and 100 forget speakers, focusing primarily on speakers seen during training, which we show ar
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית