יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SpeedrunBench: טסטבנד לבדיקת יכולות של LLM במשחקי וידאו

SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
טסטבנד לבדיקת יכולות של LLM במשחקי וידאו. SPEEDRUNBENCH מבדיקה את יכולת ה-LLM להגיע לתוצאות טובות במשחקי וידאו, כולל טסטבנד של 9 משחקים שונים.
תקציר מקורי באנגליתarXiv:2610.08076v1 Announce Type: new Abstract: Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles tha
קרא במקור המקורי