כתבה
arXiv cs.LG ·
הגנה רב-מודלית: תרגיל ספציפי להגנה על זיהוי תמונות ועץ גרפי של גישה לאינטרנט
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
מאמר זה עוסק בפיתוח תרגיל ספציפי להגנה על זיהוי תמונות ועץ גרפי של גישה לאינטרנט. התרגיל, המכונה DMAST, משתמש בשיטת שלבי תרגיל וכולל שלבי אימיטציה, הדרכה ולמידה עצמית. התרגיל נבחן באמצעות תצורה של חברה-מכונה, שבה המכונה נתקלת באטאקים שונים. התרגיל הוכיח את יעילותו בהגנה על זיהוי תמונות ועץ גרפי של גישה לאינטרנט.
תקציר מקורי באנגליתarXiv:2603.04364v2 Announce Type: replace Abstract: Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a consistent deceptive narrative. Our vulnerability analysis on MiniWob++ reveals that attacks including a visual component far outperform text-only injections, exposing critical gaps in text-centric VLM safety training. Motivated by this finding, we propose Dual-Modality Multi-Stage Adversarial Safety Training (DMAST), a framework that formalizes the agent-attacker interaction as a two-player general-sum Markov game and co-tr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית