כתבה
arXiv cs.LG ·
OpenGameEval: בנצ'מרק לתכנות אגנטי
OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
OpenGameEval הוא בנצ'מרק לפיתוח משחקים אגנטיים. הוא בודק מודלים שונים ב-84 משימות, ומודד את הצלחתם. התוצאות מראות שהמודלים הטובים ביותר מצליחים ב-51.7% מהמשימות.
תקציר מקורי באנגליתarXiv:2610.02563v1 Announce Type: new Abstract: We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4%
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית