יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MVFA: תצורה רב-מבטית לתגובה מולטימודלית לשיפור נבחנות והבנת רגשות

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition
MVFA היא תצורה רב-מבטית שמשפרת נבחנות והבנת רגשות בשיחות, על ידי שימוש במודלי LLM קפויים. התצורה נבחנה על שלושה קבצי נתונים: CH-SIMS V2.0, MELD ו-CHERMA. MVFA הציגה תוצאות שיפוט של 84.62% Acc2 ו-84.59% F1 על CH-SIMS V2.0, 67.36% Acc ו-66.03% WF1 על MELD, ו-74.66% Acc על CHERMA.
תקציר מקורי באנגליתarXiv:2609.06188v1 Announce Type: new Abstract: Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic, and visual modalities. Although large language models (LLMs) offer powerful language understanding, adapting them to multimodal affective computing remains challenging: full-model fine-tuning is computationally prohibitive, while many existing lightweight adapters fail to preserve rich textual cues during cross-modal fusion. To address these limitations, we propose the multi-view text-guided multimodal fusion adapter (MVFA), a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability. MVFA first constructs complementary text views via max pooling, mea
קרא במקור המקורי