כתבה
arXiv cs.AI ·
X-VC: תרגום דיבור זרם חסר-מטרה במרחב הקודק
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
X-VC היא מערכת תרגום דיבור זרם חסר-מטרה שמבצעת תרגום אחד-צעד במרחב הקודק של קודק עם-עצמו. המערכת משתמשת במעבד קולי דו-תנאי שמודל את הלטנטים של הקודק ואת התנאים האקוסטיים של המדבר. X-VC נבנתה על ידי צוות מחקר של Jerrister והקוד והמודלים שלה פורסמו בגיטהב.
תקציר מקורי באנגליתarXiv:2604.12456v3 Announce Type: replace-cross Abstract: Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building zero-shot VC systems for interactive scenarios remains challenging because high-fidelity speaker transfer and low-latency streaming inference are difficult to achieve simultaneously. In this work, we present X-VC, a zero-shot streaming VC system that performs one-step conversion in the latent space of a pretrained neural codec. X-VC uses a dual-conditioning acoustic converter that jointly models source codec latents and frame-level acoustic conditions derived from target reference speech, while injecting utterance-leve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית