יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קופסת טקסט-שמע ללא הצטברות לדיבוב קול וסינתזה שיחה

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
חוקרים פיתחו קופסת טקסט-שמע ללא הצטברות, המאפשרת דיבוב קול וסינתזה שיחה באיכות גבוהה. המערכת משתמשת במודל Diffusion Transformer ומאפשרת יצירה של עד דקה של שמע בפעולה אחת.
תקציר מקורי באנגליתarXiv:2609.03992v1 Announce Type: new Abstract: We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we
קרא במקור המקורי