יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

תקציר מקורי באנגליתarXiv:2610.02563v1 Announce Type: cross Abstract: We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.
קרא במקור המקורי