כתבה
arXiv cs.CL ·
MineExplorer: בחינת סוכני MLLM בעולם פתוח
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
MineExplorer הוא בנך' לבחינת יכולות חקר עולם פתוח של סוכני MLLM ב-Minecraft. הוא בוחן יכולות תפיסה, היגיון ויצירת פעולות. הבנך' משתמש בסינתזה של סוכנים רבים כדי ליצור משימות רב-שלביות.
תקציר מקורי באנגליתarXiv:2605.30931v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. Existing embodied and game-based benchmarks often compress interaction into short-horizon tasks or entangle success with domain-specific game mechanics. In this paper, we introduce MineExplorer benchmark for evaluating open-world exploration capabilities of MLLM agents in Minecraft. We first filter atomic tasks whose solutions rely heavily on Minecraft-specific knowledge to better reflect general open-world reasoning. Then we organize the benchmark around a ReAct-style capability formulation and compose atomic tasks into implicit multi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית