Nvidia hints at future AI-focused Spectrum-4 Ethernet

Applications of AI


At a time when generative AI has made machine learning a household word and has brought something of a dotcom craze to glass houses around the world, Nvidia’s data center business is putting money in its fist right now. Not surprising. It’s also no wonder that the company’s graphics card business has been sluggish, but is picking up a bit as its PC business picks up in the post-pandemic world.

Jensen Huang, co-founder and chief executive officer, found it interesting in a call between NVIDIA and Wall Street analysts to review financial results for the first quarter of fiscal year 2024, which ends in April. Chief Financial Officer Colette Kress suggested. The new Spectrum-4 Ethernet switching announcement is just around the corner. And to be a little more specific, this new networking will take place next week at the Computex trade show in Taipei, Taiwan, likely on May 30, just after the Memorial Day holiday. , Huang added. (Thank you for making all of us in the US work on the holidays.. appreciate. )

NVIDIA shares are up 23.5% on Wall Street as of Thursday morning as I write this, and the company is about to set a record for the largest single-day market capitalization in history, appearing to be the first chip company to do so. It will go down in history with a market capitalization of $1 trillion. The reason for this large and immediate surge is that the company told analysts Wednesday night after the market ended that it had sales of about $11 billion in the second quarter of the fiscal year, a give or take of about 33%. That’s about 50% higher than Wall Street expected. .

Cause? Generative AI Boom. Demand for GPU computing, not just the current ‘hopper’ H100 accelerator, has far outstripped supply, and we see no rush in pricing. the sky is so high (I’ll use the terminology), but also the previous generation “Ampere” A100 accelerators, and indeed anything AMD or Intel could manufacture if they could afford to exceed their commitments to building exascale supercomputers for national labs. also applies to By the way, they don’t have spare capacity.When GPU accelerator comparable to the cost of core We all know that gold, oil, and other very good things were mined on IBM System z mainframes.

So it’s a seller’s market, with Nvidia basically the only seller outside of the AI ​​startup-populated fringes. The good news for them is that the second half of fiscal 2024 will be even stronger than the first half, even as Nvidia ramps up supply of his A100 and H100 computing engines and his NVLink and InfiniBand networking to support them. is implying that it will become Large AI training clusters and, currently at least for large language models, AI inference clusters.

Let’s talk about networking announcements as far as we can guess, and then talk about numbers for Nvidia’s data center business. This will undoubtedly propel GPUs in the foreseeable future and will soon be the main business overall for CPUs and DPUs. .

“For multi-tenant cloud migrations that support generative AI, our high-speed Ethernet platform with BlueField-3 DPUs and Spectrum-4 Ethernet switching provides the best Ethernet network performance available,” said Kress. claimed in a call to Wall Street. “BlueField-3 is in production and has been adopted by multiple hyperscale and CSP customers, including Microsoft Azure, Oracle Cloud, CoreWeave, Baidu, and more. We look forward to sharing more about our Gb/s Spectrum-4 accelerated AI networking platform.”

In April 2022, Nvidia announced a Spectrum-4 Ethernet ASIC with an aggregate bandwidth of 51.2 Tb/s driving 128 ports running at 400 Gb/s and 64 ports running at 800 Gb/s. So it’s not a new chip, but apparently it will ship this quarter.

With Broadcom’s announcement of the ‘Jericho3-AI’ StrataDNX chip for building AI clusters on Ethernet instead of InfiniBand, Nvidia probably didn’t just blink, it jumped. Also, hyperscalers and cloud builders don’t want to deploy InfiniBand at scale in their clusters. Because in a world where these tech giants deploy his 100,000 endpoints in a single data center on Ethernet fabrics, InfiniBand endpoints reach about 40,000 endpoints. Clos topology and leaf/spine networks. Additionally, we believe NVIDIA could have charged a premium for InfiniBand given its low latency and high message rate advantages, similar to his Mellanox previous standalone. But as explained in the Jericho3-AI announcement, these large language models and recommender systems are moving large amounts of big chunky data, and message rate is less important. So instead of using these AI training workloads to adapt HPC networks based on InfiniBand to the job, hyperscalers and cloud builders are using Ethernet switch fabrics to do what AI needs, he says with Nvidia. Broadcom provides it.

Nvidia will dance on the subject and claim our interpretation is not correct – we are sure of this. But in reality, hyperscalers and cloud his builders have a near-religious enthusiasm for Ethernet, so he’s only reluctantly installing InfiniBand. Any technical discussion is welcome.

Here’s how Huang explained the situation ahead of the upcoming Spectrum-4 announcement and explained it in no time. To get the full context of when InfiniBand had record quarters and was above $1, here’s Huang’s full story: It’s $1 billion a year, and it’s going to be a record year, as the analysts who did the math a few weeks ago and analyzed such numbers denied neither Mr. Kress nor Mr. Huang. The $6.9 billion acquisition of Mellanox paid for itself long ago. Even if it was a good dream, Arm Holdings didn’t pay for his $40 billion-plus investment.

Inhale and read Mr. Huang’s quote, slightly edited. we took out hmm and ah Phrases that are repeated occasionally. It sounds natural when you speak it, but it looks weird when you print it.

“InfiniBand and Ethernet target different applications within the data center. Both have their place. InfiniBand had a record quarter. Nvidia’s Quantum InfiniBand has an extraordinary roadmap and is going to be truly incredible, but the two networks are very different: InfiniBand was designed for the AI ​​factory, if you will. If that data center runs a few applications for a small number of users in a specific use case and keeps it running continuously, and that infrastructure costs $500 million, Pick a number, the difference between InfiniBand and Ethernet can be 15% or 20% in overall throughput, and $500 million spent on infrastructure with a difference of 10% to 20%. When you get to $100 million InfiniBand is basically free that’s why people use it InfiniBand is practically free the difference in data center throughput is not negligible and one of the I’m just using it for my application, except when the data center is a cloud data center and its multi-tenant, i.e. a collection of small jobs, shared by millions of people, Ethernet is exactly the answer, there’s a new segment in the middle where the cloud is becoming a generative AI cloud, not an AI factory itself, but it’s still a multi-tenant cloud where you want to run generative AI workloads. This new segment is a great opportunity and Computex, as I mentioned in his last GTC, will be announcing a major product line for this segment, namely the Ethernet-focused generative AI application type cloud. However, InfiniBand has performed well, achieving record year-over-year numbers. “

can you see the dance Don’t get me wrong. Huang is the most eloquent conversationalist and technologist in his IT field. If Huang is dancing, it’s because the distinction is subtle. And he just fails to give the impression that InfiniBand is under any kind of pressure.

As I’ve pointed out many times before, there’s been a dance between InfiniBand and Ethernet for 25 years. And long story short, Ethernet steals all the good ideas from his InfiniBand and isn’t very good at implementing them, but keeps InfiniBand at bay and his HPC and now his AI department in the market. is enough to prevent being relegated to

But here comes the problem. Clouds and hyperscalers are eager to adopt InfiniBand to do AI research and build internal models, enabling them to run HPC and now AI training workloads on the cloud, but we believes all conditions are equal. I would prefer Ethernet like InfiniBand. And some of the big cloud networking heads said something similar recently, and we teased them a bit back in February. Because a lot of what they want Ethernet to do is what InfiniBand has been doing for a long time and has been doing well. If we had started moving to InfiniBand 20 years ago, it wouldn’t have been a big deal. But instead, InfiniBand suppliers he counts 1, Cornelis Networks he counts maybe 2, his HPC labs in China (which are different but don’t count as derived) he counts 3 down to the company.

We believe there is tremendous pressure to create leaner Ethernet that excels in HPC and AI so that networks don’t have to be different within these hyperscalers and cloud data centers. We also believe that if these pressures result in his Nvidia and Broadcom ethernet variants performing as well or better than InfiniBand in AI training tasks, this will have a significant long-term impact on the InfiniBand market. Nvidia has to respond to Broadcom’s Jericho3-AI development for over a year and Arista Networks’ general talk for over a year that they may also adopt InfiniBand for their AI training workloads. be.

We look forward to seeing what potential capabilities lurk in Nvidia’s Spectrum-4 ASIC for AI training. We’ll talk about all that next week.

So let’s spend some time with Nvidia’s numbers telling the whole story.

Nvidia’s revenue fell 13.2% to $7.19 billion in the quarter, but net income increased 26.3% to $20, thanks to a change in product mix (and a 70% gross margin on AI training gear). $40 million, representing 28% of revenue.

Nvidia’s computing and networking group saw revenue rise 21.5% to $4.46 billion, while its graphics group, which sells GPUs for PCs and workstations, fell 40.8% to $2.73 billion. .

Here’s the breakdown by Nvidia division:

The Data Center division, which sells hardware and software to Glasshouse, differs slightly from Computing and Networking, with revenue of $4.28 billion, an increase of 14.2%.

Here’s the passage that made Nvidia’s stock explode:

“I would like to move on to our outlook for the second quarter of 2024,” Kress said on the conference call. “Total revenue is expected to be $11 billion, plus or minus 2%. This demand has expanded our data center visibility over several quarters and allowed us to source significantly increased supply in the second half of the year.”

Our guess is that tech giants and other big companies will get Nvidia chips and systems, similar to what happened with many vendors at a time when supply chain tensions were at their peak during the coronavirus pandemic. I am pre-paying for it and am wondering if I should do some capacity planning with Nvidia to make sure I get it. what they need. This means NVIDIA knows more specifically about what his 2024 fiscal year will look like than he otherwise would. And Nvidia can still charge a doubled seller’s market price for that demand.

And in this market, AMD and Intel can sell any mass-produced GPU, even though it’s not compatible. And it doesn’t seem like Intel will be able to ramp up production fast enough to meet demand, and AMD is working desperately to make it happen, while at the same time calling the Instinct MI300A GPU an ‘El Capitan’. At Lawrence Livermore National Laboratory, which seems to be trying to make it compatible with supercomputers.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *