כתבה
arXiv cs.AI ·
AgentAudit: מסגרת פתוחה לבדיקת אמינות AI
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
AgentAudit היא מסגרת פתוחה לבדיקת אמינות של סוכני AI. היא בודקת את כלל התהליך של תכנון, בחירת כלים וביצוע. המחקר בדק חמישה מודלים, כולל GPT-5 ו-Claude Sonnet 5.
תקציר מקורי באנגליתarXiv:2609.09875v1 Announce Type: new Abstract: Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source. AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, combined with behavioural classification and failure attribution to pinpoint the exact stage responsible for an observed fai
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית