Google Cloud Announces New A3 Supercomputer VMs Built to Power LLM

Machine Learning


Image credit: Natalia Bulova/Getty Images

Over the past few months, we’ve seen LLM and generative AI screaming into our consciousness, but we’ve learned that training and running these models requires an enormous amount of computational power. is clear. Recognizing this, Google Cloud today announced the new A3 supercomputer virtual machine at Google I/O.

A3 is specifically built to address the considerable demands of these resource-intensive use cases.

“The A3 GPU VM features the latest CPUs, improved host memory, next-generation Nvidia GPUs, and major networking upgrades, specifically built to deliver the highest performing training for today’s ML workloads. ,” the company said in a statement.

Specifically, the company will equip these machines with Nvidia’s H100 GPUs, paired with a purpose-built data center to unleash immense computing power with high throughput and low latency. All of that prices, they suggest, are generally more reasonable than what you would pay for something like that. package.

If you’re looking for specs, 8x Nvidia H100 GPUs, 4th Gen Intel Xeon Scalable processors, 2TB of host memory, and 3.6TB of bisection bandwidth between 8x GPUs via NVSwitch and NVLink 4.0 Consider having it on board. These two Nvidia technologies are designed to maximize performance. Improve throughput across multiple GPUs like this one.

These machines can deliver up to 26 exaflops of power, helping improve the time and cost of training large-scale machine learning models. Additionally, the workloads on these VMs run on Google’s specialized Jupiter data center networking fabric, which Google describes as “26,000 highly interconnected GPUs.” This enables “full-bandwidth reconfigurable optical links that can adjust their topology on demand.” The company says this approach should also help reduce the cost of running these workloads.

The idea is to provide customers with a vast amount of features designed to train more demanding workloads, such as LLMs running complex machine learning models and generative AI applications, making it more cost-effective. It’s about doing it in a high way.

Google offers A3 in several ways. Customers can do it themselves or, if they prefer, have it done as a managed service where Google handles most of the heavy lifting. The DIY approach includes running A3 VMs on Google Kubernetes Engine (GKE) and Google Compute Engine (GCE), while managed services include running A3 VMs on Vertex AI, the company’s managed machine learning platform. To do.

The new A3 VMs will be announced today at Google I/O, but are currently only available via a preview waitlist.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *