כתבה
arXiv cs.AI ·
EvoRiskBench: בנק אבולוציוני לסיכוני ביצוע בזמן ריצה בסוכני משרד
EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents
EvoRiskBench הוא בנק אבולוציוני של סיכוני ביצוע בזמן ריצה בסוכני משרד, שמטרתו לבחון את רמת הבטיחות של סוכני משרד. הבנק כולל 450 תרגילים שונים, שנבחנו על שלושה סוכני משרד שונים: Claude, Codex ו-DeepSeek-V4-Pro-0813.
תקציר מקורי באנגליתarXiv:2610.03153v1 Announce Type: cross Abstract: Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchmark evolution. We introduce EvoRiskBench, an evolving benchmark organized around the EP-Path-EF framework, which links an initial risk entry point to a one-hop technical effect through an agent-mediated risk path. The framework defines nine entry-point categories and five effect categories; a 20-participant study supports their interpretability and classification consistency on representative cases. Guided by this framework, an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית