כתבה
arXiv cs.LG ·
Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting
תקציר מקורי באנגליתarXiv:2609.10613v1 Announce Type: cross Abstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating over competing behavioral modes, and interprets in-context demonstrations as inference-time evidence that dynamically shifts the model's posterior preference between safe and harmful behaviors. This view formalizes jailbreak as a process of evidence
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית