יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

זיהוי רגשות רב-מודאלי משופר

Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion
זיהוי רגשות רב-מודאלי הוא תחום מחקר חשוב. מחקר זה מציג גישה חדשה לזיהוי רגשות רב-מודאלי, המשלבת קידוד מרובה-תכונות ואלגוריתם פיוז'ן מבוסס תשומת לב. המודל משתמש ב-Wav2Vec2 ו-ResNet50-BiLSTM כדי לחלץ תכונות מאודיו ווידאו.
תקציר מקורי באנגליתarXiv:2609.04690v1 Announce Type: cross Abstract: Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interaction, remote education, and healthcare. This paper proposes a novel multimodal emotion recognition framework that integrates rich audio and visual feature extraction with an attention-based fusion strategy. For audio, we extract three complementary feature types: semantic embeddings from Wav2Vec2, MFCC features, and statistical acoustic descriptors such as pitch, energy, and rhythm. These are aligned and fused via a BiLSTM to capture temporal dependencies. For video, we propose a ResNet50-BiLSTM architecture that combines deep residual learning and sequential modeling to extract expressive spatiotemporal features from facial sequence
קרא במקור המקורי