כתבה
arXiv cs.CL ·
מערכת AgentActionBench: ריפודוקציה של ניסויים באמצעות סוכנים
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers
מערכת AgentActionBench נועדה לבדוק את יכולת הריפודוקציה של ניסויים באמצעות סוכנים בתחומי ML ו-AI4Science. המערכת כוללת 150 מאמרים, כולל 120 מאמרי ML ו-30 מאמרי AI4Science.
תקציר מקורי באנגליתarXiv:2609.11117v1 Announce Type: new Abstract: Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית