יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SWE-Game: יכולת הפיתוח של סוכני קוד לבניית משחקים

SWE-Game: Can Coding Agents Build the Games We Want?
אפשר לסוכני קוד לבנות משחקים שאנו רוצים? חוקרים פיתחו תקן לבדיקת יכולת הפיתוח של סוכני קוד. התקן, SWE-Game, כולל 247 משימות שמבוססות על 41 משחקים שנכתבו ב-Godot. המחברים טענו כי סוכן הקוד Opus5 הצליח לבנות משחקים שמקיימים 50.38% מהתקן.
תקציר מקורי באנגליתarXiv:2609.33678v2 Announce Type: replace Abstract: We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Gam
קרא במקור המקורי