יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

BodyCam-VQA: תיאור וידאו משופר של קמרה גוף על ידי תגובה מולטימודלית והפקת שאלות

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation
מערכת תיאור וידאו משופרת של קמרה גוף על ידי תגובה מולטימודלית והפקת שאלות. המערכת משתמשת בשיטת תגובה מולטימודלית לצפייה בווידאו ובהפקת שאלות. המערכת נועדה לספק תיאור וידאו משופר של קמרה גוף, כולל תיאור של סצנות ושל פעולות של החוקרים.
תקציר מקורי באנגליתarXiv:2609.10815v1 Announce Type: cross Abstract: Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a significant challenge due to its multimodal video format. BWC videos, in many cases, comprise chaotic scenes with low visual quality, rapid movement/interactions, and high-noise audio that make visual understanding a challenge for even SOTA multimodal models. Current Vision-Language Models (VLMs) frequently overlook critical forensic details, such as the presence of valuable evidence or the latent nuances of suspect-officer interactions, which are vital for fair legal outcomes and civilian/officer safety. To a
קרא במקור המקורי