יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Structured-Noise Masked Modeling לווידאו, אודיו ועוד

Structured-Noise Masked Modeling for Video, Audio and Beyond
אופציית הסתרה ברוטרסטרית למודלי וידאו ואודיו
תקציר מקורי באנגליתarXiv:2503.16311v2 Announce Type: replace Abstract: Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach enhances masked video and audio modeling frameworks without any additional computational cost. Experiments show that structured noise masking consistently outperforms random masking, underscoring t
קרא במקור המקורי