כתבה
arXiv cs.AI ·
אם ניקודי סוכנים עושה את מה שהם אומרים?
Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
בדיקה של חוקי חוזה ניתנים לביצוע של סביבות סוכנים שמשתמשים בכלים חשפה תכנותיות ופגמים במערכת הניקוד.
תקציר מקורי באנגליתarXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected de
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית