כתבה
arXiv cs.AI ·
Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
תקציר מקורי באנגליתarXiv:2610.03498v1 Announce Type: cross Abstract: Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-1
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית