כתבה
arXiv cs.AI ·
Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
תקציר מקורי באנגליתarXiv:2511.23304v2 Announce Type: replace Abstract: In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The task aims to mimic human reasoning by extracting and fusing information from audio-visual scenes, with the main challenge being the identification of question-relevant cues from complex audio-visual content. Existing methods fail to capture the structural information within videos and suffer from insufficient fine-grained modeling of multi-modal features. To address these issues, we are the first to introduce a new multi-modal scene graph that explicitly models objects and their relationships as a visually grounded, structured representation of the audio-visual scene, yielding 461,292 relation t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית