Post-training large language models is an engineering problem. It's not that the algorithms are hard — the formulas for RLHF/GRPO/PPO are clearly laid out in papers — it's the engineering that's hard.
You need a training framework (e.g., Megatron) to update model weights, an inference engine (e.g., SGLang/vLLM) to generate rollout data, a data buffer to manage training samples, a reward computation module to score outputs, and possibly an environment interaction module for agentic tasks. Each of these components is a complex system in its own right, and wiring them together into a complete RL training loop is the truly painful part.
slime is the open-source RL post-training framework from Tsinghua's THUDM team. You're probably more familiar with the models behind it: GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6, GLM-4.5.
This is not a "reference implementation" — it is the actual tool that trained these models.
Core Design: One Path, Three Roles
slime's most critical design decision is that training, rollout, and data generation all go through the same path.
Many RL frameworks use a "building blocks" design: the trainer is one service, the rollout engine another, data generation a third, communicating via APIs or message queues. The upside of this design is modularity; the downside is that debugging is hard — bugs in RL training are usually silent (the model appears to be learning, but it's actually learning the wrong things). When data flows through multiple services, pinpointing the source of a problem is extremely painful.
slime's design is "pipeline-style": Megatron training, SGLang rollout, custom data generation, reward computation, verifier feedback, environment interaction — all of these components flow through the same "training/rollout/data buffer" path.
The benefit of this design isn't performance (pipeline and building-block designs perform similarly), but debuggability. When all data flows through one path, you can see the complete data flow in one place: what the model generated, what score the reward function gave, what's stored in the data buffer, and which samples the trainer used to update weights.
Megatron + SGLang: Why These Two
slime chose Megatron (training) and SGLang (inference) as its core engines and exposes their parameters "natively" — not wrapped in an abstraction layer, but passed straight through.
There's a clear philosophy behind this choice: no lowest-common-denominator abstraction.
Many frameworks try to support multiple training backends (Megatron, DeepSpeed, FSDP) and multiple inference backends (vLLM, SGLang, TensorRT-LLM) simultaneously. To unify interfaces, they build a "generic trainer interface" and "generic inference interface". The problem is that different backends have different capabilities — some inference optimizations in SGLang don't exist in vLLM, and some parallelism strategies in Megatron aren't supported by DeepSpeed.
To unify the interface, a framework can only take the intersection — features every backend supports. That's the lowest-common-denominator abstraction: you sacrifice each backend's unique capabilities for compatibility.
slime doesn't do this. It only supports the single Megatron + SGLang path, but exposes the parameters of both engines directly. You can use all of SGLang's advanced features (e.g., --sglang- prefixed arguments) without waiting for the framework to adapt them.
Freedom in Data Generation: Not Just Math and Code
The core of RL post-training is training data. Different tasks require different data generation approaches:
- Math: math problems + answer verifiers
- Code: coding problems + test cases
- Search: search environments + answer matching
- Tool use: tool APIs + call validation
- Multi-agent: environments + interaction protocols
- Weight synchronization: after the trainer updates weights, the rollout engine needs to load the new weights. Mistimed synchronization leads to rollouts generated with stale weights
- Fault recovery: RL training typically runs for days to weeks; node crashes are routine. The checkpoint mechanism must restore to a correct state
- Debugging paths: slime supports rollout-only and train-only modes for debugging inference and training separately
- Reproducibility: RL training has many sources of randomness (sampling, data shuffling, distributed training); all random seeds must be explicitly recorded
- Qwen series: Qwen3.6, Qwen3.5, Qwen3Next, Qwen3MoE, Qwen3, Qwen2.5
- DeepSeek V3 series: DeepSeek V3, V3.1, DeepSeek R1
- Llama 3
- GitHub: https://github.com/THUDM/slime
- Documentation: https://thudm.github.io/slime
- GLM family: https://github.com/THUDM
Many frameworks provide a dedicated trainer for each task type — a math trainer, a code trainer, an agent trainer. The problem is that these trainers don't share infrastructure, and adding a new task type means implementing from scratch.
slime's design plugs data generation and reward computation into a unified training path as "plugins". Math, code, search, tools, sandboxes, verifiers, environments, multi-agent systems, long-horizon agentic workflows — all connect as data generation or reward workflows, without forking the training kernel.
This means adding a new task type only requires writing a data generation script and a reward function — no training code changes needed.
What "Battle-tested" Really Means
slime's README contains this line: "one of the most battle-tested open RL post-training frameworks".
The support for this claim: GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6, and GLM-4.5 were all trained with slime. This is not a framework validated on toy datasets — it was validated in real model release pipelines.
"Battle-tested" means slime has been through every engineering problem RL training throws at you:
These aren't "features" — they're "fixes born from pain". A framework that hasn't actually trained large models wouldn't think of these engineering details.
Supported Model Ecosystem
Beyond the GLM family, slime also supports:
This list shows slime isn't a GLM-exclusive tool — it's a general-purpose RL post-training framework that just happens to have been developed and validated by the GLM team.
RL Post-Training as Infrastructure
slime's emergence points to a trend: RL post-training is shifting from "research experiment" to "engineering infrastructure".
Two years ago, RLHF was a research topic requiring you to reproduce algorithms from papers, write your own training loop, and build your own rollout service. Now, frameworks like slime package the whole pipeline: you provide the model and data generation logic, and slime handles the training loop, rollout scheduling, weight synchronization, and fault recovery.
This resembles what HuggingFace Transformers did for NLP — turning large model training from a "research topic" into an "engineering task". The difference is that slime focuses not on pretraining but on RL post-training.
Once RL post-training infrastructure matures, more teams can run RL training experiments. This could give rise to a new wave of RL training methods — just as HuggingFace enabled more people to run pretraining experiments.
---
Related links: