יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SAVU-BENCH: בנק אפיון ראל-עולם להבנת ספציאל-אודיו-ויזואלי

SAVU-BENCH: A Real-World Benchmark for Spatial Audio-Visual Understanding
בנק אפיון ראל-עולם להבנת ספציאל-אודיו-ויזואלי. הבנק מבוסס על 7 תרגילים ו-3 רמות יכולת. נמצא כי רוב המודלים חסרים בהבנת קשרים ספציאליים.
תקציר מקורי באנגליתarXiv:2610.10624v1 Announce Type: cross Abstract: Spatial audio-visual understanding requires models to recognize not only what is present, but also where events occur and how they relate across modalities. Existing benchmarks often rely on simulated scenes, evaluate isolated spatial skills, and provide limited diagnostic insight into failure modes. We introduce SAVU-Bench, a real-world benchmark that systematically evaluates spatial audio-visual understanding across three capability levels and seven evaluation tasks. We further introduce SAVU-Diag, a scene-linked diagnostic set that decomposes reasoning questions into their prerequisite grounding and alignment sub-tasks. Evaluation of 12 representative models on SAVU-Bench reveals that while visual spatial grounding is relatively mature,
קרא במקור המקורי