Nvidia remains the gold standard for AI training chips, CEO Jensen Huang told investors, even as rivals try to grab the company's market share and one of Nvidia's main suppliers downplayed its AI-chip revenue outlook.
Everyone from OpenAI to Elon Musk's Tesla uses Nvidia's chips to run their large-scale languages and computer vision models, and that lead will be further strengthened when Nvidia's “Blackwell” system is introduced later this year, Huang said at the company's annual shareholder meeting on Wednesday.
Announced in March, Blackwell is the company's next-generation AI training processor, following on from its flagship “Hopper” series of H100 chips, which are among the tech industry's most expensive assets, costing tens of thousands of dollars apiece.
“The Blackwell architecture platform is likely to be the most successful product in our company's history, and indeed in the entire history of computing,” Huang said.
Nvidia briefly overtook Microsoft and Apple this month to become the world's most valuable company in a spectacular rally that has driven much of the S&P 500's gains this year. With a market capitalization of more than $3 trillion, Huang's company was once worth more than the entire economy and stock market before suffering a record decline in its value as investors booked profits.
But as long as Nvidia chips remain the benchmark for AI training, there's little reason to think the long-term outlook is uncertain, and fundamentals remain looking strong.
One of Nvidia's key advantages is its sticky AI ecosystem known as CUDA (Compute Unified Device Architecture). In the same way that average consumers are reluctant to switch from an Apple iOS device to a Samsung smartphone running Google Android, entire corps of developers have been working with CUDA for years and are so familiar with it that they have little reason to consider using a different software platform. As with hardware, CUDA has effectively become a standard of its own.
“NVIDIA's platform is broadly available across all major cloud providers and computer manufacturers, creating a large and attractive base for developers and customers, making our platform even more valuable to customers,” Huang added on Wednesday.
Micron's next quarterly earnings outlook is in line with expectations but not enough for bulls
The AI industry was hit recently when memory chip supplier Micron Technology, which supplies high-bandwidth memory (HBM) chips to companies such as Nvidia, said it expects fourth-quarter revenue to meet market expectations of only about $7.6 billion.
Micron's shares fell 7%, well below a small gain in the tech-heavy Nasdaq Composite Index.
Until now, Micron and its South Korean rivals Samsung and SK Hynix have experienced the cyclical booms and busts common to the memory chip market, which has long been considered a commodity business compared with logic chips such as graphics processors.
But excitement is building as demand for the chips needed to train AI grows. The company's stock price has more than doubled over the past 12 months, meaning investors have already priced in much of management's predicted growth.
“The guidance was basically in line with expectations, but in the world of AI hardware, it's a bit disappointing when guidance is in line with expectations,” said Gene Munster, a tech investor at Deepwater Asset Management. “Momentum investors just didn't see any additional reason to be more positive about this story.”
Analysts are closely tracking demand for high-bandwidth memory as a leading indicator for the AI industry because it is critical to solving the biggest economic constraint facing AI training today: scaling problems.
HBM chips address scaling issues in AI training
Importantly, costs do not rise with model complexity (which can number in the billions), but rather exponentially, resulting in reduced efficiency over time.
Even if revenues grow at a steady rate, losses could balloon to billions or even tens of billions of dollars a year as models become more sophisticated—a situation that could become unmanageable for a company that doesn't have an investor with the deep pockets of Microsoft to ensure OpenAI stays “solvent,” as CEO Sam Altman recently put it.
The primary reason for the declining returns is the widening gap between the two drivers of AI training performance: the raw computing power of logic chips (measured in FLOPS, or calculations per second), and the memory bandwidth needed to feed data quickly (often expressed in transfers per second, or MT/s).
Because they work in tandem, scaling one alone leads to waste and cost inefficiency, which is why FLOPS utilization, or how much compute you can actually utilize, is a key metric in determining the cost-effectiveness of your AI models.
Sold out until the end of next year
As Micron points out, data transfer rates have not been able to keep up with increases in computing power, and the resulting bottleneck (often referred to as the “memory wall”) is a major cause of today's inherent inefficiencies when scaling AI training models.
This explains why the US government was so focused on memory bandwidth when deciding which specific Nvidia chips should be banned from being exported to China in order to undermine Beijing's AI development programs.
Micron said Wednesday that its HBM business is “sold out” through the end of next calendar year, one quarter later than the fiscal year, echoing similar comments from South Korean rival SK Hynix.
“We expect to generate hundreds of millions of dollars in revenue from HBM in FY24. [billions of dollars] Micron announced on Wednesday that it expects revenue from HBM to reach $10 billion in fiscal year 2025.
