There’s no shortage of data, but there’s still a lot of video and audio data that hasn’t been used to train AI models: Google’s Jeff Dean

AI Video & Visuals


There have been some reports that companies are running out of data to train AI models, but Google’s Jeff Dean believes such concerns may be overblown.

Speaking at Nvidia’s GTC 2026 conference on March 18, Jeff Dean, principal scientist at Google DeepMind and Google Research, pushed back against the common belief that the AI ​​industry is hitting a data wall. His argument is that the world is full of untapped data on which models have yet to be trained, and the industry has only just scratched the surface.

Jeff Dean

“I also disagree in some ways with the argument that we’re running out of data,” Dean said. “I feel like there’s a huge amount of data out there that hasn’t been used to train these models yet. We’re training with some video data, but I think there’s a lot more video data with associated audio data that we’re not necessarily training on yet.”

He pointed to two particularly rich but little-tapped sources: robotics and self-driving cars. “I think the data for real-world robotics and self-driving cars will be quite rich,” he noted.

Dean also argued that synthetic data is useful as a third tool. When the interviewer asked if training on data and using that same data to generate synthetic data isn’t just producing variations of the same thing, Dean acknowledged that point, but insisted it still works. “It actually seems to be quite useful if the model is very powerful – if the model that is generating the synthetic data is very powerful.”

He went further and utilized techniques that had proven effective in previous generations of image models. “There are all kinds of techniques that were popular with convolutional image models years ago that we haven’t yet used, such as data augmentation. You can think of synthetic data as part of that. Techniques to prevent overfitting are interesting. You can use dropout or distillation as a way to normalize the model.”

The bottom line, in Dean’s view, is that the real constraint may be compute, not data. “I think there’s a lot of opportunity to make the model better by adding more compute and more passes through the data, without necessarily overfitting it.”


Dean’s comments came during a live debate. Celebrities such as former OpenAI Principal Scientist Ilya Sutskeva and Elon Musk have argued that training data is running out, with Sutskeva famously calling data the “fossil fuel of AI.” But a counter-narrative is building. Cohere CEO Aidan Gomez said the vast majority of the data the company currently uses to train new models is synthetic data. Andrej Karpathy argues that books and other existing content are not endpoints, but rather facilitators for the generation of synthetic data. And Google itself said after the release of Gemini 3 that it still benefits from pre-training, implying that it may be leveraging proprietary data that other labs don’t have access to.

This unique angle is what makes it interesting for Google. The company stores data that is simply not available to pure AI companies like OpenAI, Anthropic, and Mistral. YouTube alone hosts hundreds of millions of hours of video and audio. This is exactly the kind of multimodal data that Dean identified as underutilized. Google Maps, Google Search, Gmail, and Google’s decades of crawling the web represent a data moat that’s impossible to replicate. The company’s Gemini Robotics program is also expanding its pipeline of real-world physics data. If Dean is right that the bottleneck is not the availability of data, but the willingness and ability to use it, then Google’s exceptionally wide range of data sources could be a decisive advantage as the AI ​​race heats up. A challenger can match Google in terms of computing and talent, but replicating 25 years of data collection across the world’s most used digital products is another matter entirely.



Source link