כתבה
arXiv cs.AI ·
AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation
תקציר מקורי באנגליתarXiv:2609.06991v1 Announce Type: cross Abstract: Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a single text prompt. This capability poses new challenges for safety evaluation, as unsafe content may be conveyed through the audio track or arise only when the visual and audio tracks are interpreted jointly. Existing safety benchmarks largely focus on either generated video or generated audio in isolation and are therefore not designed to capture these risks. To close this gap, we introduce AV-SafetyBench, the first safety benchmark developed specifically for T2AV generation. AV-SafetyBench comprises a four-axis, 13-category taxonomy and 5,200 manually reviewed prompts that specify visual scenes, speech, and non-speech audio. Our ev
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית