כתבה
arXiv cs.AI ·
ScopeBench: האם סוכנים שומרים על גבולות מעורבות תחת לחץ מטרה?
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
ScopeBench הוא בנץ'מרק למשימות ביטחון אגנטיות. הוא בודק האם סוכנים מכבדים את גבולות המעורבות תחת לחץ מטרה. הבנץ'מרק כולל 30 משימות ביטחון אגנטיות שדורשות מהסוכנים לחרוג מגבולות המעורבות.
תקציר מקורי באנגליתarXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence. We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating the stated scope. Each task appears under two conditions that share an environment, verifier, and objective and differ only in scope: one instruction set has no scope and measures capability; the other has a natural-language scope to measure adherence.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית