Will the “law of scaling” where bigger is better allow AI to keep improving forever? History tells us we’re not so sure.

Machine Learning


OpenAI CEO Sam Altman, perhaps the most prominent figure in the artificial intelligence (AI) boom fueled by the launch of ChatGPT in 2022, loves the law of scaling.

These widely praised rules of thumb that link the size of an AI model to its capabilities illustrate the frenzy of the AI ​​industry, which is buying up powerful computer chips, building unimaginably large data centers, and restarting shuttered nuclear power plants.

As Altman argued in a blog post earlier this year, the idea is that an AI model’s “intelligence” is “roughly equivalent to the log of resources used to train and run it.” This means that by exponentially increasing the scale of the data and computing power involved, you can consistently generate better performance.

The large language model (LLM) scaling law, first observed in 2020 and further refined in 2022, comes from drawing lines on a chart of experimental data. For engineers, here’s a simple formula for how big to build your next model and how much performance improvement you can expect.

As AI models get bigger and bigger, will the laws of scaling continue to expand? AI companies are betting hundreds of billions of dollars to make this happen, but as history shows, things aren’t always that simple.

Scaling laws aren’t just for AI

Scaling laws can be amazing. For example, modern aerodynamics is built on them.

Engineers have discovered a way to compare miniature models in wind tunnels and test chambers to full-size planes and ships by using sophisticated mathematics called the Buckingham pi theorem to ensure that a few key numbers match.

These scaling ideas influence the design of almost any flying or floating product, not just industrial fans and pumps.

Another famous scaling idea underpinned the decades-long boom of the silicon chip revolution. Moore’s Law, the idea that the number of tiny switches called transistors on a microchip doubles approximately every two years, helped designers create today’s small and powerful computing technology.

However, there is a catch. Not all “laws of scaling” are laws of nature. Some are purely mathematical and can be held indefinitely. Others are just lines adapted to data that work beautifully Until then Too far removed from the situation for which it was measured or designed.

When the scaling law breaks down

History is littered with painful memories of broken scaling laws. A classic example is the collapse of the Tacoma Narrows Bridge in 1940.

This bridge was designed by scaling up what was working as a small bridge into something longer and slimmer. Engineers assumed that the same scaling arguments apply. That is, if a certain ratio of stiffness to bridge length worked before, it should work again.

Instead, moderate winds cause an unexpected instability called aeroelastic flutter. The bridge deck was torn apart and collapsed just four months after opening.

Similarly, the “laws” for manufacturing microchips had an expiration date. For decades, Moore’s Law (the number of transistors doubles every few years) and Dennard Scaling (many smaller transistors run faster while using the same power) have been incredibly reliable guides for chip design and industry roadmaps.

But as transistors became small enough to be measured in nanometers, those neat scaling rules began to collide with severe physical limitations.

When transistor gates were reduced to just a few atoms thick, current began to leak and they began to behave unpredictably. Additionally, the operating voltage can no longer be lowered due to background noise.

Eventually, downsizing was no longer the way to go. Chips are becoming more powerful, but now with new designs rather than just shrinking.

Is it a law of nature or a rule of thumb?

The language model scaling curve that Altman praises is real and so far has been very useful.

They told the researchers that if they gave the model enough data and computational power, the model would continue to improve. They also showed that the previous system had no fundamental limitations, just not enough resources.

But these are definitely curves that fit the data. These are different from the derived mathematical scaling laws used in aerodynamics and are more like useful rules of thumb used in microchip design. That means they are likely to be permanently non-functional.

The scaling rules of language models do not necessarily encode real-world problems, such as the limited availability of high-quality data for training or the difficulty of getting AI to handle new tasks, not to mention safety constraints and the economic difficulties of building data centers and power grids. There are no laws or theorems of nature that guarantee the eternal expansion of intelligence.

Investing in Curve

So far, the AI ​​scaling curve looks pretty smooth, but the financial curve is a different story.

Deutsche Bank recently warned of an AI “funding gap” based on Bain Capital’s estimate of a USD 800 billion mismatch between projected AI revenues and investments in chips, data centers, and power needed to sustain current growth.

JPMorgan estimates that even a 10% return on planned AI infrastructure build-outs could require approximately $650 billion in annual revenue for the broader AI sector.

We are still investigating what laws apply to Frontier LLM. Reality is likely to continue to evolve according to current scaling rules. Or new bottlenecks such as data, energy, or user willingness to pay could bend the curve.

Altman’s bet is that the LLM scaling law will continue. If so, the effects are so predictable that it might be worth building massive amounts of computing power. Meanwhile, growing anxiety on the part of banks is a reminder that some expansion talks could end up like Tacoma Narrows. In other words, beautiful curves in one context can hide nasty surprises in the next.



Source link