AI model tested on classic economic game reveals major differences from human thinking

AI News


We live in an era where large language models increasingly make choices that were once reserved for people. From composing emails to guiding business decisions, these systems shape our daily lives. This change raises pressing questions for researchers. When AI has to reason strategically, can it think like a human?

A new study tackles this problem by introducing artificial intelligence into a well-known economic experiment. The research was led by Dmitry Dagayev, Head of the Sports Research Department, Faculty of Economics, HSE University. He collaborated with colleagues Sofia Paklina and Petr Palshakov from HSE University Perm and Yulia Aleksenko from the University of Lausanne. The team worked together to test how a key AI model performed in the “Guess the Number” game, a modern version of the Keynesian beauty pageant.

The game looks simple. Each participant chooses a number between 0 and 100. The winner is the player whose numbers are closest to a fraction (often one-half or two-thirds) of the group average. However, decades of research have shown that human players rarely follow the mathematically optimal path. Instead, they reveal the limits of our reasoning, our expectations of others, and our emotional impact.

The researchers wanted to see if AI would make similar choices.

Putting AI in the shoes of humans

To find out, the team evaluated five widely used language models that will be available between 2024 and 2025. These include GPT-4o, GPT-4o Mini, Gemini-2.5-flash, Claude-Sonnet-4, and Llama-4-Maverick. Each model acted as a single participant across 16 scenarios derived from classic human experiments.

The scenarios were diverse. Some have changed the number of minutes used to calculate the winning numbers. Some people have changed the way they combine numbers in groups, using averages, medians, and maximums. Some explanations focused on the opponents themselves, such as first-year economics students, conference experts, or players described as angry or analytical.

Each model received the same instructions. You had to choose a number, explain why, and assume your opponent would match the explanation provided. All scenarios were repeated 50 times per model, with no opportunity to learn from previous rounds. This approach mirrors a one-shot experiment with human participants.

The first results were basic but important. All 4,000 responses followed the rules and fell within the 0-100 range. Almost all explanations showed some kind of strategic reasoning. It was missing in only 23 responses.

An overview of the experiments reproduced in this paper using the AI ​​player. CRT = Cognitive Reflex Test. (Credit: Economic Behavior and Organization Journal)

When AI matches human behavior and when it misses it

When the researchers compared the AI's choices to well-known human results, a clear difference emerged. In a classic experiment conducted by economist Rosemary Nagel, human players' average score was about 27 when the goal was half the group average, and almost 37 when it was two-thirds the group average. In all comparable cases, the AI ​​model chose the lower number.

Some models were very close to zero. This is the Nash equilibrium in most versions of the game. Some, such as GPT-4o, have chosen higher values, but they are still below the human average. All these differences were statistically significant.

The pattern remained even when different rules were used in the game. In the version based on the maximum number, both humans and AI chose higher values. Still, the models were different from each other. Claude Sonnet performed an average of about 35 songs, while Lama performed much fewer.

“These results show that AI reacts to changes in game structure in the same way as humans. When the theory predicts a higher or lower choice, the model moves in the same direction,” Dagaev told The Brighter Side of News.

“However, we found that an important gap had emerged. In the two-player version of the game, choosing zero was always weakly dominant; it never performed worse than any other choice. None of the models specified or explained this logic; instead, they relied on step-by-step reasoning about what the other person would do. This lack of dominant strategic thinking marks a clear difference from formal economic training,” he continued.

Replication results of Nagel (1995) using the LLM player. In Nagel (1995), 𝑛 varies from 15 to 18 from session to session. In our experiments, we fixed the number of players at 18 and assumed that the marginal effect of adding one player to a group of 15 to 18 players is low. (Credit: Economic Behavior and Organization Journal)

Differences between models and the power of scale

Not all AI behaves the same. For most pairwise comparisons, the model produced distinct means. GPT-4o and Claude Sonnet often landed in the middle. The Gemini flash varied between cautious and aggressive speculation. Llama showed the widest range.

The team investigated this further by testing Llama models of varying sizes, with parameters ranging from 1 billion to 405 billion. The pattern was impressive. Smaller models chose numbers close to typical human guesses, often near 50. Large-scale models have chosen much lower values, moving steadily toward theoretical predictions.

As the model size increased, so did the depth of inference. The larger model predicted more layers of thinking and appeared to bring selection closer to equilibrium.

Emotions, framing, and social context

The researchers also tested how sensitive the AI ​​was to context. They rewrote the prompts in different words, framed the game as a television contest, and assigned emotional states to their opponents.

Replication results of Grosskopf and Nagel (2008) using the LLM player. (Credit: Economic Behavior and Organization Journal)

The results closely reflected human behavior. Both humans and AI tended to choose higher numbers when the opponent was described as angry. Grief brought about small changes. Players described as more analytical encouraged lower than intuitive guesses.

Some models, particularly GPT-4o Mini and Llama, responded more strongly to the wording change. Yet, the overall structure of the response remained stable.

What the findings suggest

This research shows that modern AI can recognize strategic settings and adjust its behavior in predictable ways. They often behave more “rationally” than human participants by choosing lower numbers. At the same time, they often fail to identify a simple dominant strategy and assume that other strategies are more sophisticated than they actually are.

This is important beyond the laboratory. “We are currently at a stage where AI models are starting to replace humans in many tasks, enabling greater economic efficiency in business processes,” Dagaev said. However, he stressed that human behavior remains important in many decisions.

Understanding where AI works with people and where it doesn’t will determine how these systems are used in markets, policy, and daily life.

Practical implications of the research

These findings help clarify how AI operates in real-world economic environments. If a model consistently expects others to act strategically, it can misjudge the market due to emotion or limited reasoning. At the same time, the strong agreement with comparative trends suggests that they can still be a valuable tool for prediction and analysis.

For researchers, the results highlight areas where AI needs improvement, particularly in recognizing simple strategic advantages. This research provides guidance for society about when to trust AI decisions and when human judgment still matters.






Source link