כתבה
arXiv cs.CL ·
הכרה בסיכונים במודלים רב-מודאליים
Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
חוקרים גילו כי מודלים רב-מודאליים גדולים (MLLMs) עלולים להציג סיכונים בטיחותיים כאשר משלבים טקסט ותמונות. הם הציעו שיטה חדשה להפחתת סיכונים אלו.
תקציר מקורי באנגליתarXiv:2609.02082v2 Announce Type: replace-cross Abstract: Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this cross-modal safety drift and our pilot studies show that the safety response rate for such requests is substantially lower than that for requests containing explicitly unsafe text. This paper aims to systematically study this issue. First, we conduct an empirical analysis to identify representative unsafe response patterns. Building on these, we interpret model representations and attentions, revealing that visually risky cues receive limited attention and weakly trigger refusal. Motivated by the observation that safe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית