VMware + Run:ai: The compound computing platform that truly understands GenAI workloads   

AI and ML Jobs


Undoubtedly, we will remember 2023 as the year when large language models (LLMs) became the reasoning engines capable of generating both reasoning traces and task-specific actions interleaved. (see, ReAct paper). Here we are talking about closed-sourced LLMs such as ChatGPT and mainly about the plethora of open-source LLMs such as Falcon and StarCoder that are excellent for undertaking language reasoning tasks.

As LLMs get applied as core components of business applications, app dev teams must ensure they deliver the benefits of LLM capabilities to their organizations affordably, responsibly, and legally. As a result, data science and app dev teams will have to run multiple LLM instances in different phases of the app dev lifecycle, such as experimentation (e.g., instruction finetuning), integration (e.g., LangChain APIs), and deployment (e.g., LLM prompting interfaces). In such scenarios, a scalable, flexible, and cost-effective LLM operations platform will be key for success.

We are happy to announce that VMware and Run:ai have joined forces to provide a Kubernetes-based platform, leveraging VMware Tanzu, optimized to run core LLM and GenAI tasks such as finetuning and inference. This ensures you can easily operate LLMs within your organization, effectively utilize your GPU resources, and minimize GPU idle time.

Before we describe the functionality that can allow you to use dynamic and tailored partitions of any NVIDIA GPU for every job, it is important to brief you about what you can do with LLMs on a single or even a fraction of a mid-range GPU, such as the A100 (40GB).

The open source community has been constantly releasing powerful LLMs, such as Falcon-40B and Falcon-7B, which rank at the top of the Open LLM Leaderboard. It is possible to finetune these models and use them for prompt completion using a single GPU. For instance, we have published a GitHub repo that provides Python code and reproduction instructions so you can finetune Falcon-7B using a single A100 (40G) GPU and Falcon-40B on two of those GPUs by using HugginFace’s implementation of LoRA and the bits and bytes library (by Tim Dettmers) to load the models using 8-bit quantization.

Considering these LLMs will also require GPU capacity to perform fast prompt completion at inference time, you will need an infrastructure that dynamically allocates enough GPU resources to LLM finetuning and inference operations while ensuring no wasted capacity. For instance, loading the 7B parameter model using 8-bit precision requires around 17GB of GPU memory. If you run this model on an A100 (40GB), you’ll leave over 57% of that GPU unused. Wouldn’t it be great to have a way to use that spare capacity to run another instance of the same LLM for another use case? Wouldn’t it be even better to have a smart job scheduler running on top of Kubernetes that allows your MLOps and LLMOps platform to submit ML jobs for execution and leave them to the scheduler to figure out the best way to accommodate those jobs within your GPU fleet? That type of infrastructure is what VMware and Run:ai can build for you!

With Run:ai’s dynamic resource sharing and GPU fractioning features, the remaining 57% of GPU capacity can be automatically allocated to another ML practitioner’s task. This means that while one team member is finetuning a model, another can use the available resources to perform additional finetuning on a separate LLM. This way, Run:ai’s GPU fractioning accelerates the overall LLM development process.

Apart from training, the team may also choose to use the remaining 57% for inference jobs. In the inference use case, Run:ai’s auto-scaling feature allows you to increase or decrease the number of replicas dynamically. When there is high demand, Run:ai can utilize the same GPU, scaling up the number of replicas to accommodate the workload while delivering high performance. Conversely, during periods of low demand, Run:ai can autoscale the number of replicas, even down to zero, providing cost savings by minimizing resource usage. Although this approach may introduce a brief cold start delay when the first request comes in, it offers a flexible and efficient solution to adapt to varying inference demands while optimizing costs.

In addition to autoscaling, deploying models for inference with Run:ai is a streamlined process. Internal users can easily deploy their models and access them through managed URLs or user-friendly web interfaces such as Gradio and Streamlit. This enables efficient sharing and showcasing of deployed LLMs, facilitating collaboration and providing a seamless experience for stakeholders.

Sharing GPU resources among teams is not easy, as each team may have their own resources. Manual assignment of resources or under-the-desk resources to specific groups often leads to inefficiencies and idle GPU time. Run:ai addresses this by pooling all of the resources in the organization and providing guaranteed GPU quotas for teams. Users always have guaranteed access to their guaranteed GPUs without the need to monopolize resources. If idle resources exist in other teams, they can instantly utilize those idle GPUs until the guaranteed team uses them. This approach enables faster iteration and scalability of experiments, granting instant access to idle GPUs without relying on IT to request additional resources.

As an ML practitioner, you may require different types of GPUs for various workloads. For instance, using an A100 GPU would be more beneficial for training your LLM than T4s. Run:ai’s Node Pools functionality comes into play here, allowing ML teams to tailor their compute resources per job. When working with a heterogeneous cluster containing different GPU variants, Node Pools enable assigning workloads to GPUs based on their relevant compute profiles. This intelligent resource allocation ensures that each LLM task runs efficiently, optimizing performance and minimizing compute waste.

Data science preferences for experiment-tracking tools and frameworks vary widely, with countless options available. Recognizing this, Run:ai provides a seamless integration experience with popular data science frameworks and tools. With out-of-the-box integrations for PyTorch, TensorFlow, JupyterLab, Tensorboard, MLflow, Kubeflow, Ray, Weights & Biases, and many more are readily integrated, allowing data scientists to effortlessly start working with their favorite tools right away without the hassle of complex integrations or workflow dependencies. We understand that data science is already complex enough – there’s no need to complicate it further with infrastructure-related challenges.

The joined forces of VMware Tanzu and Run:ai offer a fully NVIDIA AI Enterprise certified, end-to-end, and enterprise-ready Kubernetes-based platform that integrates with your preferred tools and frameworks, streamlining your ML journey from building to training and deploying models. By leveraging unique scheduling and GPU optimization technologies, you can accelerate AI development, improve ML practitioners’ productivity, increase GPU availability, and multiply the return on your AI investment. With Run:ai’s dynamic resource sharing, you can be confident that GPU computing and memory will get efficiently allocated on-demand, freeing you from waiting for available resources.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *