abstract
Benchmarking is essential for understanding the functionality of large-scale language models (LLMs), but developing and updating real-world work tasks is costly. We use an agenttic AI approach here, where the LLM itself automatically generates and assesses practical exams for tasks across finance and business operations, management, computers and mathematics roles. To develop these exams, we distinguish between the materials needed (text, data, images, etc.) and the tools needed to solve them (function calls, web searches, etc.). Focusing on text-only tasks that do not require the use of tools, we find that only 7% (149 tasks) of these occupations are testable. To complete these comprehensive tests, we deploy 13 models, including GPT, Claude, and Gemini variants. Even for basic tasks, current LLMs struggle. The median scores of the leading models reach 65-79%, with particularly poor performance in data manipulation and financial calculations. However, the models showed rapid improvement, with models released in 2024 averaging 40.5%, while 2025 models reached 66%, an improvement of 26 percentage points in one year. Although considerable work remains to validate and extend the use of tools to extend current text-only testable tasks, our results suggest that LLM-generated benchmarks provide a cost-effective, scalable, and updatable approach to measuring AI workplace capabilities, and have the potential to extend the “LLM as judge” paradigm to the assessment of occupational tasks.
Download the full working paper
