כתבה
arXiv cs.CL ·
בטיחות רב-מודאלית רב-תורית
Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment
MINT-Safe הוא מאגר נתונים רב-מודאלי לאינטראקציות רב-תוריות. TAD-Align הוא כלי להסתגלות בטיחותית. הם משפרים את הבטיחות ב-10% ואת היעילות ב-8%.
תקציר מקורי באנגליתarXiv:2601.04736v2 Announce Type: replace Abstract: Despite remarkable capability in multi-modal understanding, deploying Multi-modal Large Language Models (MLLMs) in open-ended conversational scenarios introduces safety risks that remain poorly addressed by existing alignment methods. Unlike simple malicious visual question and answer (VQA) pairs , multi-turn interactions enable adversaries to incrementally reconstruct harmful intent across dialogues, progressively bypassing safety constraints in ways that are difficult to detect at any individual turn. Meanwhile, conventional reinforcement learning from human feedback (RLHF) approaches are unsuitable for this situation: designed primarily for VQA tasks, they neither capture cross-turn risk dynamics nor scale efficiently without costly ma
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית