כתבה
arXiv cs.AI ·
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
תקציר מקורי באנגליתarXiv:2607.16057v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית