כתבה
arXiv cs.AI ·
SceneJail: פריצה למודלי LLMs המעבדים וידאו
SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs
מאמר חדש מציג פריצה למודלי LLMs המעבדים וידאו, על ידי שימוש במצבי סצנה שונים. המאמר מציג פרקטיקה חדשה לפריצה למודלי LLMs, הקרויה SceneJail, שמשתמשת במצבי סצנה שונים כדי לפרוץ למודלי LLMs. המאמר כולל תיאור של הפרקטיקה והתוצאות של המחקר.
תקציר מקורי באנגליתarXiv:2609.38899v1 Announce Type: cross Abstract: Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding sc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית