יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

VEX-Bench: בנק אבחון לסוכנויות LLM

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities
VEX-Bench הוא בנק אבחון לסוכנויות LLM שבודקות את האפשרות לניצול פגיעויות בשרשרת האספקה. הוא כולל 75 מקרים אמיתיים מ-GitHub ומודלים כמו GPT-5.5.
תקציר מקורי באנגליתarXiv:2609.08040v1 Announce Type: cross Abstract: The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependency is actually exploitable. Security analysts typically spend substantial time assessing vulnerability exploitability case by case. Recent LLM agents have emerged as promising candidates for this task given their advanced capabilities in coding and cybersecurity, yet no existing benchmark evaluates them on it. Prior benchmarks target zero-day settings, where agents detect and exploit previously unknown vulnerabilities. In contrast, software supply c
קרא במקור המקורי