כתבה
arXiv cs.LG ·
Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection
תקציר מקורי באנגליתarXiv:2609.07455v1 Announce Type: cross Abstract: How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained? Most weakly supervised VAD methods assume centralized training; recent VLM-based extensions further rely on dense inference, generated explanations, or additional adaptation. We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. We study two VLM feedback interfaces: parsed text-generation decisions and a logit-based interface that extracts a continuous anomaly score from next-token Yes/No probabilities. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-In
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית