כתבה
arXiv cs.AI ·
VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
תקציר מקורי באנגליתarXiv:2609.09396v1 Announce Type: cross Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this "Infrastructure AI Gap." It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית