יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הצגת רעש מאורגנת: דגם מסווג לווידאו, אודיו ועוד

Structured-Noise Masked Modeling for Video, Audio and Beyond
הצגת רעש מאורגנת לדגמי ווידאו ואודיו. פיתוח חדשני ללמידת מודלים עצמית.
תקציר מקורי באנגליתarXiv:2503.16311v2 Announce Type: replace-cross Abstract: Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach enhances masked video and audio modeling frameworks without any additional computational cost. Experiments show that structured noise masking consistently outperforms random masking, undersco
קרא במקור המקורי