Openai, People's Name and Big Tech are investing billions in developing cutting-edge large language models, and a small group of AI researchers are working on the next big thing:
Computer scientists like Fei-Fei Li, the Stanford professor who invented Imagenet, and Yann Lecun, the lead AI scientist at Meta, are building what is called the “world model.”
Unlike large-scale linguistic models, which determine output based on the statistical relationships between words and phrases in training data, world models predict events based on the mental structures in which humans create the world around them.
“Language doesn't exist in nature,” Li said in a recent episode of Andreessen Horowitz's A16Z podcast. “Man,” she said, “We not only survive, live and work, but we build civilization beyond language.”
In his 1971 paper, “Counsel Behavior of Social Systems,” computer scientist and MIT professor, Jay Wright Forrester explained why mental models are important for human behavior.
We use each model all the time. Every person in personal life and business instinctively uses models for decision-making. The mental imagery in your mind about your surroundings is a model. Your head doesn't include any actual family, business, city, government, or country. Use selected concepts and relationships to represent the actual system. Mental image is a model. All decisions are based on the model. All laws are passed on the model. All executive actions are based on the model. The problem is not to use or ignore the model. The problem is only choice between alternative models.
If AI meets or outweighs human intelligence, the researchers behind it believe that it should also be able to create mental models.
Li is working on this through World Labs, which she co-founded in 2024, with initial support of $230 million from venture companies such as Andreessen Horowitz, New Enterprise Associates and Radical Ventures. “We aim to lift AI models from pixels in 2D planes to 3D world full (both virtual and real).
In the NO Priors podcast, Li stated that spatial intelligence is “the ability to understand, reason, interact and generate the 3D world.”
Li said he is looking at the application of world models in creative fields, robotics, or any field that justifies the infinite universe. Like Meta, Andrill and other Silicon Valley heavyweights, it could mean advances in military applications by helping battlefield people better perceive their surroundings and predicting the next move of the enemy.
The challenge of building a world model is lack of sufficient data. In contrast to the language that humans have refined and documented over the centuries, spatial intelligence is less developed.
“If you ask them to close their eyes now and pull out or build 3D models of the environment around you, that's not that easy,” she said on the No Priors podcast. “Until you're trained, you don't have that much ability to generate very complex models.”
Collecting the data needed for these models “requires increasingly sophisticated data engineering, data collection, data processing, and data synthesis,” she said.
It furthers the challenge of making a world of faith even bigger.
In Meta, AI scientist Chief Yann Lecun has a small team dedicated to similar projects. The team uses video data to train models and perform simulations that abstract the video at different levels.
“The basic idea is not to predict at the pixel level. You can train your system to perform an abstract representation of a video and make predictions with that abstract representation.
This creates a simpler set of building blocks to map trajectories about how the world changes at a particular time.
Lecun, like Li, believes these models are the only way to create truly intelligent AI.
“We need an AI system that allows us to learn new tasks very quickly,” he recently said at the National University of Singapore. “They need to understand not only text and language but the physical world, the real world. They have some level of common sense, ability to reason and plan, and have a lasting memory.
