כתבה
arXiv cs.AI ·
Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
תקציר מקורי באנגליתarXiv:2610.01833v1 Announce Type: new Abstract: Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing al
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית