יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ShieldCLIP: הסרת תוכן מזיק במודלים רב-מודאליים

ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
ShieldCLIP הוא כלי להסרת תוכן מזיק ממודלים רב-מודאליים. הוא משתמש במידע על בטיחות כל מודל כדי להסיר תוכן מזיק. הכלי נבדק על מודלים כמו CLIP ו-Stable Diffusion.
תקציר מקורי באנגליתarXiv:2609.39688v1 Announce Type: cross Abstract: Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dat
קרא במקור המקורי