יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למידת יציבות זמן-תדר לצורך זיהוי דיקליים דיבור

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
במאמר זה, נחקרת יציבות זיהוי דיקליים דיבור בסצנריות ריאליות. המחברים מציגים פרקטיקה ללמידת יציבות זמן-תדר, שמטרתה ללמוד תיאורי דיקליים דיבור שישמרו תכונותיהם גם לאחר עיבוד קו השמע. הם ניתחו את השפעת עיבוד קו השמע על יציבות זיהוי דיקליים דיבור והציגו פרקטיקה לשיפור יציבות זיהוי דיקליים דיבור. הקוד זמין ב-https://github.com/JunXue-tech/TFCL.
תקציר מקורי באנגליתarXiv:2607.17761v2 Announce Type: replace-cross Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Ti
קרא במקור המקורי