כתבה
arXiv cs.AI ·
A2Z GameSpec-Bench: כמה נאמן יכולים סוכני קוד לייצר משחקים מפרטי תכנון של משחק?
A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
במאמר זה, נוצרה בסיסת-תקן חדשה לבדיקת יכולתם של סוכני AI לייצר משחקים מפרטי תכנון של משחק. הבסיסת-תקן, A2Z GameSpec-Bench, כולל 100 פרטי תכנון של משחקים ארוכים, ובודקת את יכולתם של הסוכנים לייצר משחקים שמקיימים את דרישות התכנון. הבסיסת-תקן נועד לבדוק את יכולתם של הסוכנים לייצר משחקים שמקיימים את דרישות התכנון, ולא רק לייצר משחקים שהם 'נאמנים' באופן כללי.
תקציר מקורי באנגליתarXiv:2609.39564v1 Announce Type: new Abstract: Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית