כתבה
arXiv cs.LG ·
WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents
תקציר מקורי באנגליתarXiv:2605.23220v2 Announce Type: replace Abstract: Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated automated evaluation methods. A key obstacle is that attack evaluation must be both accurate and efficient: weak manually tuned attacks can overestimate robustness, while exhaustive hyperparameter search is prohibitively expensive because each candidate requires closed-loop rollouts through learned latent dynamics. We introduce WMAttack, an automated attack-search framework for adversarial evaluation of world-model agents. WMAttack formulates robustness evaluation as a finite-budget search over attack configurations, including attack families, perturbation budgets, optimization steps, restarts,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית