Google’s GKE Labs has introduced OpenRL, an open source project that provides a self-hosted API for post-training and fine-tuning large-scale language models (LLMs) on standard Kubernetes clusters.
Google says OpenRL abstracts reinforcement learning (RL) infrastructure from AI research, allowing machine learning teams to scale post-training workflows directly on their own clusters.
According to Google engineers, when using agent reinforcement learning in LLM, it’s “very easy to get bogged down in system complexity.” Even a single RL loop requires juggling many moving parts, including data preparation and cleaning, environment selection, training loop debugging, reward design, inference mismatch handling, hardware provisioning, and underlying infrastructure management.
All of these are difficult questions. But what makes it even more complicated is how closely AI research and infrastructure concerns are interwoven in today’s tools and frameworks.
Google engineers argue that by separating infrastructure from AI research, these challenges become more manageable and specialized teams can focus on their areas, just as Kubernetes enables infrastructure abstraction and simplifies workflows for application developers and reliability engineers.

One way to make post-training fine-tuning more efficient in OpenRL is to run multiple RL jobs on your infrastructure to increase overall GPU utilization. According to Google researchers, because traditional RL loops are strictly sequential, the GPU often sits idle while waiting for CPU- or network-dependent tasks to complete, especially for reward calculations.
Additionally, Google points out that OpenRL improves the user experience by clearly separating responsibilities. Researchers can focus on developing the RL loop, and engineers can run and scale post-training fine-tuning workflows.
If you’re doing research and development, you don’t need to run your RL loop directly on a machine with a GPU. All you need to do is run an RL loop on your Mac pointing to the training API running on your Kubernetes cluster/VM.
The OpenRL repository contains autoresearch A recipe that shows how to run parallel experiments with parameter sweeps and adjust reward signals in a text-to-sql workflow for Gemma models. Beyond practical applications, Google highlights this as an example of how automation can streamline and scale AI research.
OpenRL is easy to use on macOS, Nvidia GPUs, and GKE. It also integrates with Tinker-Cookbook thanks to the Tinker-compatible endpoint.
OpenRL is not the only effort focused on simplifying post-training fine-tuning by better separating concerns. For example, FeynRL ensures separation of fine-tuning recipes and system logic, making it easier for researchers to develop and test new methods while extending those approaches with tools such as DeepSpeed, Ray, and vLLM.
