כתבה
arXiv cs.AI ·
הפיכת וידאו למוזיקה למשחקי וידאו
Video-to-Music Generation for Gameplay Videos
במאמר זה נחקרה האפשרות להפיכת וידאו למוזיקה בתחום משחקי הווידאו. נפתחה נתונים חדשים של 217.6 שעות של וידאו משחקי SNES עם 485 שעות של רצועות קלט. נטוען דגם פשוט של מעבד-מעבד טרנספורמר שמעביר תכונות וידאו ישירות למעבד MusicGen. נבדקו שלושה דרכי קודד: תיאורים טקסטואליים (T5), רצועות independent (ViT), או רצועות spatiotemporal (ViViT).
תקציר מקורי באנגליתarXiv:2609.31810v2 Announce Type: replace-cross Abstract: Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of clean soundtracks, free of sound effects and voice-overs, matched to gameplay audio via audio fingerprinting. With this dataset, we train a simple encoder-decoder transformer that passes video features directly to a MusicGen decoder, comparin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית