יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LogiScope-VQA: בדיקת מודלים רב-מודאליים לזיהוי סיכונים בתעשייה

LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios
LogiScope-VQA הוא מאגר נתונים לבדיקת מודלים רב-מודאליים לזיהוי סיכונים במערכות תעשייתיות. המאגר כולל 2,476 תמונות ו-2,918 סרטונים מפארקים לוגיסטיים, ו-10,274 שאילתות וידאו מאומתות. הניסויים הראו פער משמעותי בין ביצועי המודלים, כולל GPT-5.5, Gemini-3.1-Pro ו-Claude-Opus-4.7, לבין ביצועי בני אדם.
תקציר מקורי באנגליתarXiv:2609.09790v1 Announce Type: cross Abstract: Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks align
קרא במקור המקורי