כתבה
arXiv cs.AI ·
ממצאים ממקורות קומפוננטליים לפסק-דין בבטיחות: חיפוש סיבתי במודלי שפה-תמונה
From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models
במאמר זה נחקרים מודלי שפה-תמונה כדי לגלות מתי החלטות בבטיחות נעשות. נמצא כי חיפוש סיבתי במודלי שפה-תמונה חשופים את המיקום שבו החלטות בבטיחות נעשות.
תקציר מקורי באנגליתarXiv:2610.07514v1 Announce Type: new Abstract: A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other exa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית