כתבה
arXiv cs.AI ·
EnterpriseBench: ניטענות תקן לבדיקת כושר של סוגי סוכני LLM בתחום ההחלטה האסטרטגית בעסקים
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
ניטענות תקן לבדיקת כושר של סוגי סוכני LLM בתחום ההחלטה האסטרטגית בעסקים. התקן, הנקרא EnterpriseBench, כולל שלושה סביבות פעילות: Consulting, Beer Game ו-Enterprise Digital Twin.
תקציר מקורי באנגליתarXiv:2609.37658v1 Announce Type: new Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introd
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית