כתבה
arXiv cs.AI ·
התקפת Spatial Temporal Coherence Adversarial נגד מודלים של Vision Language
Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
חוקרים פיתחו התקפת Spatial Temporal Coherence Adversarial נגד מודלים של Vision Language המשמשים בנהיגה אוטונומית. ההתקפה מצליחה לרמות את המודלים Qwen2.5-VL-7B ו-Video LLaVA-7B.
תקציר מקורי באנגליתarXiv:2610.08331v1 Announce Type: cross Abstract: The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית