כתבה
arXiv cs.LG ·
Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
מקדמים טכנולוגיה לאופן קביעת אירועים ויזואליים-קוליים עם פורטפוליו פתוח
תקציר מקורי באנגליתarXiv:2610.11846v1 Announce Type: cross Abstract: Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית