יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

KaliBench: תקן ניטור דק-משטח לכלי תגבור צבאי בקאלי לינוקס

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
KaliBench הוא תקן ניטור דק-משטח לכלי תגבור צבאי בקאלי לינוקס, המאפשר ניטור יותר דק של כלי תגבור צבאי. התקן כולל 8,504 זוגות פקודה-פקודה, המכסים 1,642 כלים ב-23 תחומי יכולת ו-5 שלבי ביטחון. KaliBench נבנה דרך פייפלינג עם קנוניזציה תקינה ובדיקה חכמה, המאפשרת ניטור זהיר וניתן להעתקה של תקינת כלי תגבור צבאי.
תקציר מקורי באנגליתarXiv:2610.02206v1 Announce Type: cross Abstract: LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability d
קרא במקור המקורי