There are growing concerns among some investors that the AI sector, which has single-handedly kept the economy out of recession, has become an unsustainable bubble. Nvidia, a major supplier of chips used in AI, became the first company to be worth $5 trillion. Meanwhile, OpenAI, the developer of ChatGPT, is yet to turn a profit and burns through billions of dollars of annual investment. Still, financiers and venture capitalists continue to pour money into OpenAI, Anthropic, and other AI startups. Their bet is that AI will transform every sector of the economy and that jobs will be replaced by technology, just as happened with typists and switchboard operators in the past.
But there's also reason to fear this gamble won't pay off. For the past 30 years, AI research has been organized around improving narrowly specified tasks, such as speech recognition. However, with the advent of large-scale language models (LLMs) like ChatGPT and Claude, AI agents are increasingly being asked to perform tasks without a clear way to measure improvement.
For example, consider the seemingly mundane task of creating a PowerPoint presentation. What makes a good presentation? We may be able to point to best practices, but the “ideal” slideshow relies on the creative process, expert judgment, pacing, sense of narrative, and subjective taste, all of which are highly contextual. An annual review presentation is different from a startup pitch or project update. You know a good presentation when you see it, but you know it's a bad presentation when it fails. However, the standardized tests currently used in the field to evaluate AI cannot capture these qualities.
This may seem like a small problem, but reputation crises have contributed to historic AI failures. And without being able to accurately measure how good the AI actually is, it's hard to know whether we're heading toward a different AI.
read more: AI architect named TIME's 2025 People of the Year
The birth of AI is often traced back to a small workshop in Dartmouth in 1956. The workshop brought together computer scientists, psychologists, and others with a common interest in mimicking human intelligence in machines. The field quickly found a powerful backer in the Defense Advanced Research Projects Agency (DARPA), an agency within the Department of Defense tasked with maintaining technological superiority during the Cold War. To keep up in the scientific race, DARPA has generously provided large, unconditional grants to AI researchers at universities and private companies over the next 40 years.
The field's first decades were defined by peaks of excitement when new technologies were invented, followed by troughs of disappointment when they failed to develop into useful applications. In the 1980s, this cycle was accelerated by AI technologies called “expert systems,” which promised to build machines with the intelligence of experts such as doctors and financial planners. Within these programs, human expertise was incorporated into formal rules such as testing patients for measles if they had a fever and rash.
Expert systems attracted significant attention and investment from the industry based on early successes such as automating loan applications. However, this optimism was fueled primarily by hype rather than rigorous testing. In reality, these expert systems tended to make strange, and sometimes disastrous, mistakes when challenged with more complex tasks. In one humorous show, the expert system suggested that the man's infection may have been caused by a previous amniocentesis, a procedure performed on pregnant women. It turns out the researchers forgot to add a rule regarding gender.
At the time, avid AI critic Hubert Dreyfuss described these failures as the “first-step fallacy,” arguing that equating expert systems with progress toward true intelligence is “like claiming that the first primate to climb a tree took the first steps toward flying to the moon.” The problem was that as the task became more complex, the number of rules needed for every possible case grew exponentially. Just like going from tic-tac-toe to checkers to chess, the number of possibilities doesn't just increase, it explodes exponentially.
When it became clear that expert systems could make no further progress, AI research entered the so-called “AI winter” in the late 1980s. Subsidies dried up, businesses closed, and AI became a dirty word.
In the aftermath, DARPA reevaluated its AI funding strategy. Rather than awarding unconditional grants, government program managers began making awards conditional on achieving the highest scores on standardized tests they called “benchmarks.” In contrast to complex problems like medical diagnostics, benchmarks focused on bite-sized tasks that are achievable and have immediate commercial and military value. We also used quantitative metrics to validate our results. Can your system accurately translate this text from Russian to English, transcribe this audio fragment, or digitize the text in these documents? Researchers didn't just make fancy claims based on promising but imperfect technology. To receive funding, companies had to provide concrete evidence of benchmark improvement.
These benchmarking competitions unified an otherwise unruly field by focusing AI researchers on common problems. Rather than each research group choosing its own project, DARPA shaped the field's collective agenda by funding researchers to work on specific tasks, such as digit recognition or speech-to-text. The competitive nature of the new funding regime crowded out AI orientations that had less success in benchmarking. For example, the first benchmark competition demonstrated that “machine learning” algorithms that can learn from data dominate the hand-crafted rule-based approaches of the past.
Soon, public leaderboards were built that provided real-time feedback on which algorithms currently held the highest scores on each benchmark, allowing researchers to learn from past successes. Once a task was resolved, a more complex task was placed in its place. Translating words translated paragraphs and eventually translated multiple languages. Number recognition gave way to object recognition in images, and then came video.
In the early 2010s, benchmarks convinced researchers to go all-in on one machine learning approach inspired by the human brain, called artificial neural networks or “deep learning,” and advances have accelerated since then, underpinning today's generative AI. Within a few years, speech-to-text algorithms powered modern AI assistants, and tumor recognition algorithms began to outperform radiologists for some cancers. The benchmark seemed to be the first step toward usable AI in everyday life.
By the end of the decade, the field was surprised to discover that advances in benchmarking tasks had led to the development of deep learning algorithms that could generate fluent and socially appropriate texts such as screenplays and poems. These abilities are do not have They appear in the benchmark because the benchmark is not designed to find them. This revelation was the catalyst for the generative AI revolution, leading to the large language models that dominate the market today, such as ChatGPT and Claude. It was the biggest victory in this field. However, with this new technology, the field faces new crises.
Simply put, there is no longer a clear benchmark for the tasks we are currently trying to automate. There is no “correct” PowerPoint, marketing campaign, scientific hypothesis, or poem. Unlike object recognition, where there are right or wrong answers, these are complex, creative, multidimensional, process-based problems, and even the most difficult benchmarks cannot objectively measure progress.

As a result, new models from ChatGPT, Claude, Gemini, and Copilot are now evaluated as much by “vibe tests” as by concrete benchmarks. We are currently caught between two inappropriate approaches. One is old-style benchmarking, which accurately measures narrow capabilities, and the other is qualitative evaluation, which attempts to understand the practical capabilities of these systems, but does not provide clear quantitative evidence of progress. Researchers are exploring new evaluation systems that bridge these perspectives, but this is a very difficult problem.
Current investments assume significant automation within the next 3-5 years. But without reliable evaluation methods, we won't know whether LLM-based technologies are leading us toward true automation or repeating the Dreyfus fallacy and taking the first step down a dead-end path. This is the difference between future infrastructure and bubbles. At this point, it's hard to tell which one we're building.
Bernard Koch is an assistant professor of sociology at the University of Chicago who studies how evaluation shapes science, technology, and culture. David Peterson is an assistant professor of sociology at Purdue University who studies how AI will transform science.
Made by History guides readers beyond the headlines with articles written and edited by expert historians. Learn more about Made by History at TIME. Opinions expressed do not necessarily reflect the views of TIME editors.
OpenAI and TIME have a license and technology agreement that allows OpenAI to access TIME's archives.
