How Containers, LLMs, and GPUs Fit for Data Apps

AI and ML Jobs


Containers, Large Language Models (LLMs), and GPUs provide the foundation for developers to build what Nvidia CEO Jensen Huang describes as an “AI factory.”

Huang made the statement at the start of the Snowflake Summit in Las Vegas this week. There, he demonstrated the power of containers as a new foundation for generative AI distribution through application architectures that integrate with enterprise data stores, with Nvidia and Snowflake.

Announced at the conference and now in private beta, Snowflake’s Snowpark Container Service (SCS) provides the ability to containerize LLM. Through Snowpark, Snowflake provides Nvidia’s GPUs and his NeMO, Nvidia’s “end-to-end cloud-native enterprise framework” for building, customizing, and deploying generative AI models.

Snowflake’s investment in container services shows how the company has transformed from its roots as a data warehouse services provider to a data analytics company and now a platform for container services and native application development. increase.

The Snowpark serves as the home of the SCS. It provides a platform for deploying and processing Python, Java, and Scala code on Snowflake through a set of libraries and runtimes, including user-defined functions (UDFs) and stored procedures.

SCS gives developers the ability to manage and scale containerized workloads (jobs, services and service functions) using a secure Snowflake management infrastructure with configurable hardware options such as GPUs. Provides developers with additional Snowpark runtime options. ”

Hex is an early user of the service. CEO and co-founder Barry McCardell Describe Hex as Figma for data or Google Docs for data. The service will allow data analysts and data scientists to work on his one collaborative platform.

“We use it (SCS) to deploy software, and now Hex deploys on Amazon Web Services, so this is interesting,” he said. “We are planning to release GCP later this year. But the environment that I am really most excited about is Snowflake. Being able to run Hex workloads where your customer data is without any additional processes is amazing. because I think it’s very attractive.”

The business case for SCS comes down to compliance and governance, which people like McCardell cite as Snowflake’s core values. SCS enables companies like Hex to offer their customers a way to build applications based on Snowflake data.

“Trying to do business with them outside of Snowflake Container Services could take months, years of security reviews and InfoSec work because many of these big customers are very cautious. It’s about enabling third-party applications to connect to data,” McCardell said. “So for the customer, it’s actually going to be very seamless. With minimal effort and minimal overhead, he’ll be able to use Hex on top of his existing data store.”

With SCS, Snowflake has the potential to provide services that alleviate many of the operational burdens that slow data synchronization for teams. Prior to the introduction of SCS, managed data managed by Snowflake was moved off the platform and containerized, creating an administrative burden.

SCS is built on Kubernetes and works behind the scenes for you. A unique take on the Snowflake environment. A customer can import containers into his SCS with little operational overhead compared to before.

Torsten Grabs, senior product manager at Snowflake, said: “Customers often cause data to be taken out of Snowflake and put elsewhere, and they also introduce redundant copies of the data somewhere in cloud storage, for example. It happened often,” he said. “Then we would run containerized compute on that data to do something with it, and sometimes the results would come back to Snowflake. , it becomes very difficult to maintain them over time and keep them in sync.What is the truth for you?Is data governance managed in a consistent way across all these places? It’s a lot easier if you can bring the work that the container is doing to where the data is and apply processing to the data there, which is one of the key motivations for us to bring containers to the platform. was.”

And here is the point. Customers often need to orchestrate containers elsewhere where GPUs meet their computing needs, such as AI and machine learning workloads. So if a user wanted a computation by his GPU, the data had to be stored in his Snowflake and retrieved to a location where the computation could be performed.

Snowflake is built on a similar approach to providing data warehouse services when the company was founded.

“When you create one of these services, you can specify which instances you want to run,” said Christian Kleinerman, Senior Vice President of Products at Snowflake, in a conversation at the Snowflake Summit. “It’s a unique mapping of logical instances narrowed down. So we can say high memory, low memory, GPU, non-GPU, which he maps to the appropriate instance for each of the three clouds his provider. mapping.”

SCS opens up opportunities to use underlying models in Snowflake.

“One way is to bring your own model or take different models and run containers,” said Kleinermann. “The heavy lifting is either you do it or you use a third party. You can do the configuration and publish it as a native app. And now I don’t know anything about AI or ML or containers, and I’m using the underlying model.”

Nvidia connection

NeMO has two main components, Kleinermann said. It comes with a specific model trained by Nvidia. It is provided as a whole framework including API and user interface. This is useful for training a model from scratch or fine-tuning a model using data fed into it. The NeMO framework is hosted within his SCS. NeMo itself is provided as a container, allowing porting of models to SCS.

Models may be imported and built on top of NeMO. For example, Snowflake announced her Reka as a partner. Her just-launched Reka creates a generative model. AI21Labs is also a Foundation Model Partner.

Kari Anne Briski is VP of AI Software at Nvidia. Briski said Nvidia is ahead of almost everyone in model development. Snowflake took advantage of the Snowflake Summit to announce that developers will be using her Nvidia GPUs and their trained models to build generative AI applications. Briski said Snowflake customers are likely to use a large foundation model she builds on Nvidia’s products.

Briski tracks his work at Nvidia as an AI development timeline, showing how Snowflake benefits from Nvidia’s research. Seven years ago, Nvidia accelerated computer vision with a single GPU. Nvidia currently uses thousands of GPUs to train its underlying models.

Briskey said training the underlying model still takes weeks or months. Briskie said that providing a pre-trained model greatly reduces the compute used.

Teams can customize models at runtime using “zero-shot or few-shot” learning, which provides a way to get answers with very little provided data, he said.

“You can send prompts, some examples, and run-time prompts that you can customize so that you can say, ‘Oh, I see what you’re talking about.’ “Well, I’m going to follow your lead.”

Dozens or hundreds of examples are available with options for Prompted Tuning or Parameter Efficient Fine Tuning (PEFT).

“We train a small model that a large language model uses, so we get this customized model,” says Briski. “We can have hundreds and thousands of customized models.”

According to Hugging Face’s blog post, PEFT significantly reduces computational and storage costs by allowing the user to fine-tune only a few (additional) model parameters and freeze most parameters in the pre-trained LLM. means to reduce. ”

All weights across the network may also change, but that comes with more intensive computing requirements.

Using LLM alone can make it vulnerable to hallucinations, and this is the case with vector databases.

“Again, don’t think of LLM in isolation,” says Briski. “You might think of it as a system as a whole. We also do these tweaks.

But overall, the concept of containers, LLMs, and GPUs means faster capabilities, more robust products, and the realization that we can interact with data ushering in a new era, says Hearthstone with Snowflake CEO Frank Slootman. During the conversation, Huang said.

“We will all be intelligence makers in the future,” Huang said. “Of course you hire employees and then you build a bunch of agents. We will connect AIs, and we will operate these AIs at scale, and we will continuously improve these AIs, and we will all manufacture AIs, You will be running an AI factory.”

Disclosure: Snowflake covered the airfare and hotel costs for the reporter’s attendance at the Snowflake Summit.

group Created in sketch.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *