כתבה
arXiv cs.AI ·
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
תקציר מקורי באנגליתarXiv:2602.12304v5 Announce Type: replace-cross Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a reference image $I^{r}$ and a reference audio $A^{r}$, this novel task requires generating videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. To this end, we propose OmniCustom, a powerful DiT-based audio-video
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית