יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התקפות על דגמי תמונה-לשון: אופטימיזציה משולבת של פקסל והזמנה

Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
מחברים תקפות על דגמי תמונה-לשון על ידי אופטימיזציה משולבת של פקסל והזמנה. הם פיתחו תקפה חדשה שמשפרת את יעילות ההתקפה על דגמי Qwen2.5-VL-7B ו-BLIP-2.
תקציר מקורי באנגליתarXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prom
קרא במקור המקורי