From rules and inference to learning from data

Machine Learning


The evolution of artificial intelligence (AI) has been extensive, from simple rule-based systems to today's highly sophisticated generative models. Through decades of breakthroughs, a distinctive part of this story has been the relentlessly accelerating thirst for data, a trend that has grown not only in scale but also in strategic importance. Modern foundational models operate on vast and ever-expanding datasets, creating new “laws” of progress. To remain competitive, AI companies must continually double the size of their training data every year or risk falling behind.

The earliest forms of AI were rule-based or expert systems, developed in the mid-20th century to mimic human reasoning through preset logical structures. Models such as General Problem Solver (GPS) and MYCIN behaved as if they were following “if-then” instructions. Although this structure allowed it to solve very narrow problems precisely, it was fundamentally rigid and suffered from scalability problems. As the problem space grew, the number of rules required became unmanageable.

As the digital revolution accelerated data creation and digital storage in the 1990s, researchers sought new approaches. Enter machine learning (ML), a paradigm shift. Rather than encoding every rule, ML systems can ingest data, learn statistical relationships, and dynamically improve performance over time. This change has been driven not only by new algorithms, but also by the wealth of new data available, a boon that has changed the way we approach AI development.

The rise of data-driven AI models

Machine learning meant that models could learn from examples, not just human instructions. Early applications used relatively modest datasets for tasks such as email spam filtering, customer segmentation, and optical character recognition. As the capacity of digital storage has increased and the Internet has become more widespread, access to larger datasets has increased both the sophistication and predictive accuracy of algorithms.

In the 2000s, advances in statistical inference, neural networks, and support vector machines pushed the boundaries. The models have become more adaptable, but the real breakthrough has come by leveraging vast amounts of new data such as web content, sensor measurements, and social media streams. These resources have enabled “big data” AI, where performance scales with the amount of data with few visible limits.

Deep learning and the scale revolution

The 2010s ushered in deep learning, neural networks with many layers that can represent increasingly abstract and complex relationships. These systems, enabled by better hardware, cloud computing, and open datasets, have revolutionized areas such as image classification, speech recognition, and natural language processing.

At the core of this revolution were convolutional neural networks (CNNs) for computer vision and transformer architectures for languages. This is exemplified in the 2017 paper “Attending Is All You Need,” which introduced transformers and set the stage for explosive growth in model scale. The impact of deep learning goes beyond technical achievements. AI has reshaped itself as a field where increasingly large data and model sizes generate new capabilities, often circumventing the need for fundamental algorithmic breakthroughs.

Scaling methods and the data arms race

A surprising discovery in recent years is that scaling up data, model parameters, and computing resources predictably improves performance across a variety of tasks. These “laws of scaling” have guided the architecture and strategy of every major AI company. Large-scale language models such as GPT-2 and GPT-4 currently rely on training datasets counted in billions or even trillions of tokens (blocks of text or data that allow the model to learn patterns, relationships, and nuances).

For example, GPT-2 (2019) was trained with approximately 4 billion tokens. By 2023, GPT-4 required nearly 13 trillion tokens. This shows how rapidly demand for data is growing. Today's state-of-the-art foundational models routinely use datasets thousands of times larger than the entire English Wikipedia, ushering in a new era in scale.

Foundation Model Era: Data as a Lifeline

Fundamental models, large-scale neural networks with multimodal understanding and generative creativity, are now the basis for a wide range of applications. These models are “pre-trained” on vast and diverse datasets and “fine-tuned” for specific tasks, domains, and industries. These are the engines behind conversational AI, autonomous vehicle recognition, generative art, and more.

All recent analyzes point to a core truth. This means that the performance, generalization, and emergent ability of such models are closely related to the size and diversity of the training data. Companies that invest in the acquisition, curation, and management of increasingly large data sets will be able to unlock new capabilities such as complex inference, multi-step planning, and creative synthesis, capabilities that only occur at scales never before seen.

Large-scale computing and parameters: The scale triad

Data is one factor, but as datasets have grown, model size (measured in trillions of parameters) and computing resources (measured in petaFLOPs) have followed a similar exponential trajectory. Modern AI training often requires weeks or months of distributed computation across tens of thousands of GPUs or specialized hardware.

Between 1950 and 2010, the compute used to train AI models doubled approximately every two years. Since 2010, that rate has doubled every six months. Currently, the largest models require training investments in the tens of millions of dollars and are only accessible to deep-pocketed organizations and frontier-focused multinational corporations.

Data curation, diversity and quality

A key frontier for building larger and more functional models is collecting relevant, high-quality, and diverse training data. Data curation and filtering have become important because low signal or repetitive data can hinder model training or lead to undesirable outputs. The underlying model team employs heuristics, automatic filters, and sampling techniques to maximize data signals and relevance. This technique becomes increasingly important as data volumes explode.

Synthetic data generation and augmentation (artificially creating new training samples) allows companies to go beyond the limitations of existing human-generated data. However, the study warns that recursively training AI-generated material can lead to diminishing returns or poor results (the so-called “model collapse” problem).

Competitive imperative: double your data or fall behind

Perhaps the most striking lesson of the past decade in AI is that continuously scaling data is essential. Companies that don't double their training data every year will quickly be overtaken by competitors leveraging larger, richer datasets. Data, computing, and exponential growth in model size go hand in hand. Slack in any area not only delays improvement but also leads to loss of capacity and market opportunities.

This is supported by empirical evidence. As your data scales logarithmically, benchmark performance, emergent inference skills, fact reproduction, and robustness improve at a predictable rate. The most ambitious players in the industry are pursuing the continuous acquisition of new sources such as text, images, video, code, and sensor streams, in some cases complementing synthetic expansion and search mechanisms.

Lack of data: a looming plateau

A provocative prospect is the depletion of high-quality human-produced training materials. At the current pace, some researchers estimate that the world's supply of useful text, images, and audio could be completely consumed within a decade. This will propel the field toward new frontiers, including simulated data, artificial environments, higher-fidelity generative processes, and innovative curation mechanisms.

The challenges are significant. Researchers warn that as AI models are increasingly trained on their own output, they run the risk of loss of diversity, propagation of bias, and recursive regression. Investing in broader and deeper data sources, such as multilingual content, scientific literature, and human interactions, remains a strategic necessity.

The future of AI and data

Looking to the future, the fate of AI is deeply intertwined with the fate of data. The era of scaling is likely to last as long as we benefit from larger datasets and smarter curation. As hardware costs fall, cloud platforms become more popular, and infrastructure improves, even mid-tier players will be able to take advantage of vast amounts of training at a lower cost.

While research continues into better algorithms, more efficient architectures, and alternative learning paradigms, the laws of scaling suggest that “more data” will remain the dominant lever for years to come. Monitoring, predicting, and understanding the impact of growing datasets is critical not only for technological competitiveness but also for aligning AI advances with ethical, social, and regulatory priorities.

Conclusion: Data is the key to the AI ​​era

From the early rule-based automata to the transformative decade of deep learning and the generative models that are reshaping today's markets, a hunger for data has been central. The constant need for more and better data drives all innovations, competitive advantages, and frontier capabilities in AI. As foundational models and their successors evolve, this desire for data will continue to be a hallmark, driving companies to new sources, creative expansion, and more sophisticated approaches to curation and diversity.

In the coming years, success in AI will mean not only more intelligent algorithms, but also smarter, larger, and higher-quality datasets. The evolution of AI has led to a simple lesson: those who feed their models the most information thrive the most.



Source link