כתבה
arXiv cs.AI ·
OmniReasoning: הגבלות הסקירה המשותפת של אודיו-ויזואלי
OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
אומני-רזונינג: הגבלות הסקירה המשותפת של אודיו-ויזואלי. חברת arXiv:2609.39490v1. המאמר עוסק בפיתוח בנק אלפא, מנוע נתונים, ושיטת הלמדה. המאמר גם מציג מודל OmniReasoning-30B-A3B שהשיג 50.0% על OmniVideoBench ו-42.5% על OmniReasoningBench.
תקציר מקורי באנגליתarXiv:2609.39490v1 Announce Type: cross Abstract: Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-v
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית