כתבה
arXiv cs.LG ·
ReH-FUSE: פרקטיקה יציבה להתאמה של מומחים לזיהוי הרגשות המודלי בשיחה
ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation
ReH-FUSE היא פרקטיקה יציבה שמשלבת מומחים שונים לזיהוי הרגשות המודלי בשיחה. היא משתמשת בטכנולוגיות שונות, כולל טקסט, אודיו ומודלים קרוס-מודליים. ReH-FUSE היא פרקטיקה יציבה שמסוגלת להתאמה למצבים שונים ולזיהוי הרגשות המודלי בשיחה.
תקציר מקורי באנגליתarXiv:2609.13857v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (ERC) requires adapting to the instance-dependent reliability of different evidence sources. Lexical content may be decisive, vocal expression may provide complementary cues, or accurate recognition may require cross-modal interaction; fixed fusion does not explicitly account for this variation. We propose ReH-FUSE, a reliability-aware framework with dialogue-aware text, audio, and cross-modal experts. Its decision-level router first models the relative preference between text and audio and then balances the resulting unimodal mixture against the cross-modal expert. This factorization separates unimodal competition from cross-modal selection. Across three independent runs on IEMOCAP, ReH-FUSE ach
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית