Nvidia unveils inference microservices to help deploy AI applications in minutes

Applications of AI



Nvidia CEO Jensen Huang, giving a keynote address at the Computex trade show in Taiwan, spoke about using Nvidia NIM (Nvidia Inference Microservices) to transform AI models so that AI applications can be deployed in minutes instead of weeks.

He said 28 million developers around the world can now download Nvidia NIM, an inference microservice that delivers models as optimized containers and deploys them to the cloud, data centers and workstations. This makes it easier to build generative AI applications like co-pilots and chatbots in minutes instead of weeks, he said.

These new generative AI applications are becoming increasingly complex and often utilize multiple models with different capabilities for generating text, images, video, voice, etc. Nvidia NIM dramatically improves developer productivity by providing a simple and standardized way to add generative AI to applications.

NIM also allows businesses to maximize their infrastructure investments. For example, running Meta Llama 3-8B on NIM will generate up to 3x more generated AI tokens on faster infrastructure than without NIM. This allows businesses to increase efficiency and generate more responses using the same amount of computing infrastructure.


Lil Snack & Game Beat

GamesBeat is excited to partner with Lil Snack to bring you customized games just for our viewers. As gamers ourselves, we know this is an exciting way for you to get involved by playing the GamesBeat content you already love. Play the games now!


Nearly 200 technology partners, including Cadence, Cloudera, Cohesity, DataStax, NetApp, Scale AI, and Synopsys, have integrated NIM into their platforms to accelerate the adoption of generative AI for domain-specific applications such as co-pilots, code assistants, digital human avatars, etc. Hugging Face is currently offering NIM starting with Meta Llama 3.

“Every company wants to incorporate generative AI into their operations, but not every company has a dedicated team of AI researchers,” Huang said. “NVIDIA NIM is helping the technology industry because it's integrated into any platform, accessible to any developer and runs anywhere.
We are putting generative AI within the reach of every organization.”

Enterprises can deploy AI applications using NIM into production through the Nvidia AI Enterprise software platform, and starting next month, members of the Nvidia Developer Program will have free access to NIM for research, development and testing on their preferred infrastructure.

40+ Microservices Power Gen AI Models

NIM is useful for a variety of businesses, including healthcare.

NIM containers are pre-built to speed up model deployment for GPU-accelerated inference and can include Nvidia CUDA software, Nvidia Triton inference server, and Nvidia TensorRT-LLM software.

At ai.nvidia.com, you can experience over 40 Nvidia and community models as NIM endpoints, including Databricks DBRX, Google's open model Gemma, Meta Llama 3, Microsoft Phi-3, Mistral Large, Mixtral 8x22B, and Snowflake Arctic.

Developers can now access the Nvidia NIM microservice for the Meta Llama 3 model from the Hugging Face AI Platform, making it easy for developers to access and run Llama 3 NIM with just a few clicks using NVIDIA GPU-powered Hugging Face inference endpoints on their cloud of choice.

Companies can use NIM to run applications that generate text, images, video, voice and digital humans. The Nvidia BioNeMo NIM microservices for digital biology enable researchers to build new protein structures to accelerate drug discovery.

Dozens of healthcare companies have deployed NIM to power generative AI inference across a range of applications, including surgical planning, digital assistants, drug discovery, and clinical trial optimization.

Hundreds of AI ecosystem partners incorporate NIM

Platform providers such as Canonical, Red Hat, Nutanix and VMware (acquired by Broadcom) support NIM in their open source KServe or enterprise solutions, and AI application companies Hippocratic AI, Glean, Kinetica and Redis have also adopted NIM to power generative AI inference.

Leading AI tools and MLOps partners such as Amazon SageMaker, Microsoft Azure AI, Dataiku, DataRobot, deepset, Domino Data Lab, LangChain, Llama Index, Replicate, Run.ai, Securiti AI, and Weights & Biases are also incorporating NIM into their platforms to enable developers to build and deploy domain-specific generative AI applications with optimized inference.

Global system integrators and service delivery partners Accenture, Deloitte, Infosys, Latentview, Quantiphi, SoftServe, TCS and Wipro have built a NIM competency to help enterprises globally rapidly develop and deploy AI strategies.

Enterprises can run NIM-enabled applications virtually anywhere, including on Nvidia-certified systems from global infrastructure manufacturers such as Cisco, Dell Technologies, Hewlett Packard Enterprise, Lenovo and Supermicro, and server manufacturers such as ASRock Rack, Asus, Gigabyte, Ingrasys, Inventec, Pegatron, QCT, Wistron and Wiwin. NIM microservices are also integrated with Amazon.
Web Services, Google Cloud, Azure, Oracle Cloud Infrastructure.

Industry leaders include Foxconn, Pegatron, Amdocs, Lowe's and ServiceNow.
Manufacturing, medical,
Financial services, retail, customer service, etc.

Foxconn, the world’s largest electronics manufacturer, is using NIM to develop domain-specific LLMs embedded into various internal systems and processes for AI factories for smart manufacturing, smart cities, and smart electric vehicles.

Developers can try Nvidia microservices for free at ai.nvidia.com. Enterprises can deploy production-grade NIM microservices with Nvidia AI Enterprise running on Nvidia-certified systems and major cloud platforms. Starting next month, members of the Nvidia Developer Program will have free access to NIM for research and testing.

Nvidia Certified Systems Program

Nvidia has certified its systems.

With the help of generative AI, companies around the world are building “AI factories” that input data and generate intelligence.

And by making its technology a critical must-have, Nvidia is helping companies deploy validated systems and reference architectures to reduce the risk and time it takes to deploy specialized infrastructure capable of supporting complex, compute-intensive generative AI workloads.

Nvidia ALSO today announced the expansion of its Nvidia Certified Systems program to designate leading partner systems as suitable for AI and accelerated computing, enabling customers to deploy these platforms with confidence from the data center to the edge.

Two new certification types have been added: Nvidia Certified Spectrum-X Ready Systems for AI in the data center and Nvidia Certified IGX Systems for AI at the Edge. Each Nvidia Certified system has undergone rigorous testing to be verified to deliver enterprise-grade performance, manageability, security and scalability for Nvidia AI.

Enterprise software workloads including generative AI applications built with Nvidia NIM (Nvidia Inference Microservices). The system provides a trusted path for designing and implementing efficient and reliable infrastructure.

The Nvidia Spectrum-X AI Ethernet platform, the world's first Ethernet fabric built for AI, combines the Nvidia Spectrum-4 SN5000 Ethernet switch series, Nvidia BlueField-3 SuperNICs and network acceleration software to deliver 1.6x AI network performance over traditional Ethernet fabrics.

Nvidia-certified Spectrum-X Ready servers serve as the building blocks for high-performance AI computing clusters, supporting the powerful Nvidia Hopper architecture and Nvidia L40S GPUs.

Nvidia Certified IGX Systems

Nvidia specializes in AI.
Nvidia specializes in AI.

Nvidia IGX Orin is an enterprise-ready AI platform for industrial edge and medical applications, featuring industrial-grade hardware, a production-grade software stack and long-term enterprise support.

It includes the latest technologies and built-in enhancements for device security, remote provisioning and management, delivering high-performance AI and proactive safety for low-latency, real-time applications in areas such as medical diagnostics, manufacturing, industrial robotics and agriculture.

Key Nvidia ecosystem partners are expected to receive the new certification: Asus, Dell Technologies, Gigabyte, Hewlett Packard Enterprise, Ingrasys, Lenovo, QCT and Supermicro will offer certified systems shortly.

Additionally, certified IGX systems will soon be available from Adlink, Advantech, Aetina, Ahead, Cosmo Intelligent Medical Devices (a division of Cosmo Pharmaceuticals), Dedicated Computing, Leadtek, Onyx and Yuan.

Nvidia also said it will become easier than ever to deploy generative AI within the enterprise: Nvidia NIM, a set of generative AI inference microservices, will work with KServe, an open-source software that automates running AI models at scale for cloud computing applications.

This combination will enable generative AI to be deployed like any other large-scale enterprise application.
Red Hat.

The integration of NIM into KServe extends Nvidia's technology to the open source community, ecosystem partners and customers. Through NIM, all of these users can access the performance, support and security of the Nvidia AI Enterprise software platform with an API call – the push button of modern programming.

Meanwhile, Huang said Meta’s openly available, state-of-the-art, large-scale language model, Meta Llama 3 (trained and optimized using Nvidia accelerated computing), is helping to dramatically improve healthcare and life sciences workflows, delivering applications aimed at improving patients’ lives.

Now available as an Nvidia NIM inference microservice available for download at ai.nvidia.com, Llama 3 gives medical developers, researchers and enterprises the means to innovate responsibly across a wide range of applications. NIM comes with a standard application programming interface that can be deployed anywhere.

With use cases ranging from surgical planning and digital assistants to drug discovery and clinical trial optimization, Llama 3 enables developers to easily deploy optimized generative AI models for co-pilots, chatbots and more.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *