High-performance computing options are dwindling, and HPE is one of the last US companies to develop supercomputers. This fact was indirectly acknowledged by University of Utah Professor Danreed in his ISC conference keynote last month.
HPE now has new compute options for running HPC-focused machine learning applications without building and managing on-premises supercomputers. The company is expanding its GreenLake cloud service to add high-performance computing options for large language models.
HPE GreenLake for Large Language Models in the Cloud says, “A single large-scale AI and HPC job can run on hundreds or thousands of CPUs or GPUs simultaneously. It’s very different from general-purpose cloud offerings: “It’s a single instance,” Justin Hotard, executive vice president and general manager of HPC Labs and AI, said at a press conference.
Supercomputers require their own data center capabilities for power and cooling, and “previously supercomputers weren’t available on demand in a consumption model,” Hotar said.
Cloud supercomputing options are part of a bigger announcement for HPE to enter the AI cloud market. HPE GreenLake for Large Language Models provides access to supercomputers, software stacks, and services for customers to run machine learning applications remotely. GreenLake packages computing as a utility service, much like a utility bills monthly. In this case, HPE charges for AI computing consumed in the cloud and provides the ability to customize software and hardware connectivity.
HPE specifically provides access to the Cray XD supercomputer, and “the first large-scale language model is on the Nvidia H100 GPU,” Hotard said, adding that “when launching additional specific types of instances, We will provide more details,” he added. The company last year unveiled the XD2000 and XD6500 supercomputers for enterprise and AI applications. The supercomputer uses some of his Cray technology, which is also used in exascale systems. The XD supercomputer features a SlingShot interconnect and a ClusterStor E1000 storage system.
An ongoing debate in the supercomputing community revolves around the relative safety of air-gapped supercomputing systems and the risk and performance issues of moving HPC workloads to the cloud. It’s being done step by step. With LLM, the cloud can be a bottleneck because it doesn’t provide the throughput or bandwidth of on-premises supercomputers. HPE is exploring how on-premises and cloud models of AI can complement each other.
“We have been running a small development cloud testbed over the last few months and have gotten a lot of positive feedback on requirements. I think they are,” Hotard said, adding that HPE GreenLake for LLM in the cloud is a complementary service to on-premises supercomputers. Customers running on-premises AI workloads want bursting capabilities and can add cloud options to their clusters. The cloud frees up on-premises resources to run more critical AI workloads and also adds diversity to HPC software and computing stacks, Hotard said.
Services and software
HPE has a stack of both levels of hardware and software services that complement each other. The software stack has features that make LLM reliable and accurate. Machine learning data management software keeps your data visible, consolidated, tracked, and audited. These features provide guardrails for generating secure and trusted data. This was a major concern that LLM would hallucinate and produce erratic responses.
“Our supercomputer leverages the HPE Cray programming environment, giving developers the tools to write or debug code, tune and optimize their applications,” said Hotard.
Users will have access to Luminous, Aleph Alpha’s large-scale language model with 13 billion parameters. LLM is multimodal. That means it can handle images and text. Top AI companies such as Google are building multimodal LLMs that support any kind of input, including voice. “This is the first service we will announce at HPE GreenLake, and we plan to release services in other areas such as climate modeling, drug discovery, financial services, manufacturing and transportation,” he said. Luminous fits in with HPE’s plans to provide a base model to support HPC workloads, not the general-purpose workloads used for general-purpose computing. HPE will initially offer an open source model and a proprietary model available for purchase.
“For example, we have already had deep discussions on the pharmaceutical side. These will change depending on the use case and which partners we work with,” Evan Sparks, chief product officer for artificial intelligence at HPE, said in response to a question. said. HPC wire.
HPE currently has no plans to offer OpenAI’s GPT-4 large-scale language models to supercomputing customers, but “that could evolve as the partnership discussions evolve,” Sparks said. .
HPE GreenLake for LLM will be generally available in North America first, followed by Europe later this year or early next year.
competition
Indeed, public cloud providers also offer supercomputing options via virtual machine instances. Google announced last month its A3 supercomputer with 26,000 Nvidia H100 GPUs. Amazon also offers its own suite of HPC products, including his EC2 instances with fast Elastic Fabric Adapter interconnects and Luster or ZFS file systems. Amazon also talks about confidential computing in its HPC products to keep data safe while it’s being processed in VMs.
Customers have multiple choices for AI and supercomputing in the cloud, depending on whether the organization uses HPE’s services, which are hybrid systems and less publicly available, or public cloud providers with security mechanisms. increase. Protect your data.
But HPE said GreenLake for LLM is complementary to its public cloud service, allowing customers to store structured and unstructured data used to train machine learning models.
“Data is the key input for training and tuning this kind of model. [mechanism] Get data into your service from the public cloud or wherever your data resides. Those mechanisms will become available,” Sparks said. GreenLake customers can also make her API calls. This is a popular way of connecting AI applications to well-known LLMs such as his GPT-4.
