יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SWE-Test: ניקוי חולשות LLM דרך חיזוי קלט

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
SWE-Test: ניקוי חולשות LLM דרך חיזוי קלט. המאמר עוסק בפיתוח תקן לבדיקת חולשות LLM ובהשוואת תוצאות של מודלים שונים.
תקציר מקורי באנגליתarXiv:2609.06229v1 Announce Type: cross Abstract: Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall against an unknowable vulnerability set, often rely on synthetic bugs, and report a single end-to-end verdict that cannot localize where an agent fails. Vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and iteratively correct from feedback. We recast its measurement as an input-prediction task with a closed, deterministic ground truth: using coverage-gui
קרא במקור המקורי