Research preview. OpenRL is an early-stage project from GKE Labs. Expect the API surface and architecture to keep evolving.
OpenRL implements Tinker compatible API for fine-tuning language models that you can run on your own infrastructure (machine or a kubernetes cluster). You can use the Tinker SDK to orchestrate RL training loops by writing imperative Python code directly from your local machine.
📖 For the full story behind why we built OpenRL, read our introductory blog post: Introducing OpenRL: A self-hosted post-training API for fine-tuning LLMs (Google Open Source Blog, June 2026).
Agentic RL on LLMs carries a lot of systems complexity. Running a single RL loop means coordinating dataset selection and cleaning, choosing RL environments, debugging the training loop, managing reward signals, handling inference mismatches, allocating hardware, and operating the infrastructure underneath all of it.
Our view is that AI research and infrastructure concerns are too tightly coupled in today's tooling. Separating them lets infrastructure engineers and AI researchers move independently — much the way Kubernetes separated infrastructure concerns from application development.
That is why we built on Tinker. Tinker simplifies LLM post-training for developers and researchers by hiding post-training infrastructure behind four API primitives. Researchers keep complete control over their training algorithms, data loops, and loss functions; platform engineers scale, orchestrate, and operate the infrastructure independently. OpenRL lets you run those same APIs on hardware you control.
Bonus: you can use tinker-cookbook that has awesome tutorials/recipes and utilities!
A traditional RL loop runs sequentially: trainers wait on samplers, and samplers wait on environments to score rewards — work that is often CPU- or network-bound. Expensive GPUs sit idle for much of the loop. Because the API decouples the loop from the hardware, OpenRL can run multiple RL jobs concurrently and pack their training and sampling steps together, which lifts overall GPU utilization.
Putting infrastructure behind an API also means researchers stop wrestling with heavy Python and CUDA dependency stacks. During R&D you can run your RL loop on a Mac while pointing it at training APIs hosted on a Kubernetes cluster or a pool of VMs, then scale up without rewriting the loop.
We expect frontier AI research to become increasingly automated, and abstracting away the infrastructure is groundwork for that. The autoresearch recipes, adapted from Karpathy's autoresearch, run parallel experiments for parameter sweeps and reward-signal improvement against a shared OpenRL gateway.
- It is not a managed service. OpenRL is self-hosted. The goal is that it is easy to deploy and operate on your own Kubernetes cluster.
- It is not an RL framework. You keep full control over your RL loop; OpenRL only provides the training and sampling APIs underneath it.
- Follow the Pig Latin notebook or Text-to-SQL notebook to see supervised fine-tuning in action.
- Follow the Text-to-SQL RL recipe to see reinforcement learning in action.
Snippet below shows a sample Reinforcement Learning loop like GRPO, where the 4 API primitives are used to create a generate-and-reward-train loop:
import asyncio
import tinker
from tinker import types
# Placeholder Environment & Reward Functions
def generate_math_problem() -> str: ...
def compute_advantages(rewards: list[float]) -> list[float]: ...
def parse_and_score_response(text: str) -> float: ...
async def rlvr_loop():
service_client = tinker.ServiceClient(base_url="http://localhost:8000")
# 1. Create Model
training_client = await service_client.create_lora_training_client_async(base_model="Qwen/Qwen3-4B-Instruct-2507", rank=16)
for epoch in range(10):
# 2A. Extract sampling client from current weights
sampling_client = training_client.save_weights_and_get_sampling_client(name=f"rlvr_epoch_{epoch}")
prompt_text = generate_math_problem()
# 2B. Sample multiple rollouts (e.g. N=8) from the prompt
response = sampling_client.sample(
prompt=types.ModelInput.from_ints(tokens=[...]), num_samples=8, sampling_params=types.SamplingParams(max_tokens=100, temperature=0.9)
).result()
# 3. Score the rollouts using the environment
rewards = []
for seq in response.sequences:
text = decode(seq.tokens)
rewards.append(parse_and_score_response(text))
advantages = compute_advantages(rewards)
# ... package sequences, text, and advantages into datums ...
# 4. Forward-Backward Pass (Importance Sampling)
# We pass the advantages to RL objective function
await training_client.forward_backward_async(datums, loss_fn="importance_sampling", loss_fn_config={"clip_range": 0.2})
# 5. Optimizer Step
await training_client.optim_step_async(types.AdamParams(learning_rate=1e-5))
asyncio.run(rlvr_loop())Detailed guides and runnable examples are structured under docs/ and examples/:
- Guides:
- Supervised finetuning:
- Reinforcement Learning:
- Technical Documentation:
- Deployment:
The FY 2026 roadmap lives in ROADMAP.md. It covers reliability and production
readiness, multi-GPU support, TPUs, more models, observability, recipes, and community. To
propose a change, open an issue with the roadmap label.
We are grateful to the open source AI communities whose work inspired OpenRL, in particular Thinking Machines, vLLM, PyTorch, prime-rl, verl, SkyRL, and llm-d.
We welcome contributions! Please see CONTRIBUTING.md for how to set up a development environment, find an issue to work on, and open a pull request.
Participation in this project is governed by our Code of Conduct.
- Questions and discussion: GitHub Issues
- Reporting a vulnerability: see our Security Policy — please do not open a public issue for security reports
- How the project is run: GOVERNANCE.md
- Who maintains it: MAINTAINERS.md
This project is licensed under the Apache 2.0 License.
This is not an officially supported Google product.
This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.