כתבה
arXiv cs.CL ·
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
תקציר מקורי באנגליתarXiv:2506.04141v2 Announce Type: replace-cross Abstract: The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to match frames mentioned in the question (hereafter referred to as "question frame") and perceive a few adjacent frames. To address this gap, we propose MMR-V: A Benchmark for Multimodal Deep Reasoning in Videos. The benchmark is characterized by the following features. (1) Long-range, multi-frame reasoning: Models are required to infer and analyze evidence frames that may be far from the question frame. (2) Beyond perception: Questions cannot be answered throug
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית