כתבה
arXiv cs.CL ·
אופטימיזציה של סגנון עוין: שיפור התקפות Jailbreaks
Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization
חוקרים גילו כי מודלים רב-מודאליים רגישים להתקפות Jailbreaks. הם הציגו אופטימיזציה של סגנון עוין (ASO) כדי לשפר התקפות אלו. ASO משתמשת ב-GRPO כדי לאפטים תמונות עוינות.
תקציר מקורי באנגליתarXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying performance against the rapidly evolving MLLMs, failing to exploit non-content-based vulnerabilities. Unlike previous research, we empirically find that MLLMs exhibit a Stylistic Inconsistency between their comprehension ability and safety ability: MLLMs can robustly understand content regardless of visual style, yet their defense mechanisms can be easily bypassed by specific stylistic triggers. Based on this finding, we propose Adversarial Style Optimization (ASO), a plug-and-play enhancement module to amplify existing
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית