כתבה
arXiv cs.AI ·
ERPbench: בדיקת סוכני LLM להחלטות עסקיות בתחום השוק התחרותי
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
ERPbench - בדיקת סוכני LLM להחלטות עסקיות בתחום השוק התחרותי. המחקר מציג בדיקה של 6 מודלי LLM, כולל DeepSeek ו-Gemini, בתחום השוק התחרותי.
תקציר מקורי באנגליתarXiv:2609.04667v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-l
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית