כתבה
arXiv cs.AI ·
AVIO: הוספה והסרה של אובייקטים משמיעים בסצנות אודיו-ויזואליות
AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
AVIO היא שיטה להוספה והסרה של אובייקטים משמיעים בסצנות אודיו-ויזואליות. היא משתמשת במודל קודם-אימון ליצירת אודיו-ויזואלי ומתאימה אותו לצורך הוספה והסרה של אובייקטים. AVIO מאפשרת עריכה של סצנות אודיו-ויזואליות באופן יעיל ומדויק.
תקציר מקורי באנגליתarXiv:2609.36503v1 Announce Type: new Abstract: Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית