כתבה
arXiv cs.AI ·
חקירת טרנספורמטורים של דיפוזיה לצורך אוגמנטציה חצי-תעבודתית בקידוד מצבי מודלי-מודלי
Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding
במאמר זה, המחברים מציגים טכניקה חדשה לקידוד מצבי מודלי-מודלי, המשתמשת בטרנספורמטורים של דיפוזיה. הם מציעים אוגמנטציה חצי-תעבודתית של מודלי-מודלי, המאפשרת קידוד טוב יותר של מצבי מודלי-מודלי. המחברים מציגים תוצאות של ניסויים שהראו תוצאות טובות יותר מאשר טכניקות אחרות.
תקציר מקורי באנגליתarXiv:2609.11341v1 Announce Type: new Abstract: Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propose CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation that treats paired modalities as sources of mutual generative supervision rather than merely as inputs to be fused. CoMA-DiT conditions velocity prediction on the paired modality through cross-modal attention and adaptively injects the resulting variation via a reliability-gated residual mechanism. Experiments on multimodal auditory attention decoding and emotion recognition showed that CoMA
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית