יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

תקציר מקורי באנגליתarXiv:2608.11469v2 Announce Type: replace-cross Abstract: AI agents are rapidly improving in cybersecurity when source code is available, yet much of the software most consequential to security, including malware, firmware, and proprietary applications, exists only as binaries. Analyzing such software requires reverse engineering (RE): recovering program semantics before analysis can proceed. Evaluating agentic RE poses a fundamental challenge: realistic benchmark instances must (1) be absent from LLMs' training data to prevent shortcuts by memorization, and (2) reflect the scale and anti-analysis protections of real-world binaries. We introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built from scratch by RE experts with over 5,000 expert hours, SRE-Bench comprise
קרא במקור המקורי