כתבה
arXiv cs.AI ·
CheckerBench: האם סוכני זמן-ארוך יכולים לייצר נקודות-בדיקה סטטיות?
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
במאמר זה, נוצרה בנקדת ניסויים (CheckerBench) של 300 תרגילים, שנבנו מ-297 CVEs ו-167 רפרטואריות. הבנקדת נועדה לבדוק את יכולתם של סוכני זמן-ארוך לפתח נקודות-בדיקה סטטיות.
תקציר מקורי באנגליתarXiv:2610.07557v1 Announce Type: cross Abstract: Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that inde
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית