כתבה
arXiv cs.CL ·
Vision-Language Models: כיצד לשמור על עזרה כאשר פוגעים באופן בלתי מובהן?
Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment
מודלי Vision-Language נותרים חשופים לסכנות בלתי מובהנות. חידוש חדש משפר את הבטיחות והעזרה. החידוש, OPSD, משתמש בהנחיות רצונות כמידע פרטי במהלך הלמידה, כדי לעזור למודלים לזהות סכנות בלתי מובהן ולספק תשובות בטוחות ומועילות. OPSD מציג תוצאות טובות יותר בהשוואה לשיטות בטיחותיות קיימות, ומציע פתרון יעיל יותר לבעיית הבטיחות במודלי Vision-Language.
תקציר מקורי באנגליתarXiv:2609.37837v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית