כתבה
arXiv cs.AI ·
Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation
תקציר מקורי באנגליתarXiv:2607.23493v1 Announce Type: cross Abstract: Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית