כתבה
arXiv cs.AI ·
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
תקציר מקורי באנגליתarXiv:2605.10834v3 Announce Type: replace Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as capture-the-flag, remote code execution, exploit reproduction, or trajectory similarity, in simplified or narrow settings. These tools are valuable for measuring bounded capabilities, yet they do not adequately capture the complexity, open-ended exploration, and strategic decision-making required in realistic pentesting. In this paper, we present a practical evaluation protocol that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית