TRL

Fine-tune and align language models with SFT, DPO, GRPO, and RLHF trainers built on Hugging Face Transformers.

Library
PyPI
v1.14.0
19,403 stars
Apache License 2.0

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
93 /100 Excellent
Development Activity 96
Maintenance 100
Community 76
Maturity 60
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
88 /100 Excellent
Architecture 85
Code Quality 90
Innovation 85
Learning Curve 90

TRL (Transformer Reinforcement Learning) is Hugging Face’s library for post-training foundation models. It ships dedicated trainer classes — SFTTrainer, DPOTrainer, GRPOTrainer, KTOTrainer, RewardTrainer, RLOOTrainer, and a DistillationTrainer — each a thin, well-tested wrapper around the Transformers Trainer that adds the loss functions, data collation, and generation loops specific to its post-training method. Because every trainer shares a common base class and config pattern, moving from supervised fine-tuning to preference optimization to RL-style policy training is a matter of swapping the trainer and config classes rather than re-deriving the training loop.

The library is built directly on top of the Hugging Face ecosystem: Transformers for models and tokenizers, Datasets for data loading, Accelerate for distributed training (DDP, FSDP, DeepSpeed ZeRO), and it integrates with PEFT for LoRA/QLoRA training on constrained hardware and with vLLM for fast generation during RL-style training loops like GRPO. A trl CLI wraps the most common recipes (trl sft, trl dpo, trl kto) so straightforward fine-tuning jobs can be launched without writing a training script, while a trl.experimental namespace holds newer, less stable trainers (async GRPO, distillation variants, and more) that graduate into the stable API as they mature.

What You Get

  • Trainer classes for the major post-training methods — SFTTrainer, DPOTrainer, GRPOTrainer, KTOTrainer, RewardTrainer, and RLOOTrainer — each paired with a matching Config dataclass.
  • A trl command-line tool (trl sft, trl dpo, trl kto) for launching standard fine-tuning and preference-optimization jobs without writing training code.
  • Native integration with Accelerate for multi-GPU/multi-node training (DDP, FSDP, DeepSpeed ZeRO) and with PEFT for LoRA/QLoRA on modest hardware.
  • A reward-function library (trl.rewards) with ready-made verifiers (e.g. accuracy_reward, reasoning_accuracy_reward) for GRPO-style RL training on reasoning tasks.
  • An experimental namespace (trl.experimental) exposing newer trainers — async GRPO, distillation variants, replay-buffer GRPO — ahead of their stabilization.

Common Use Cases

  • Supervised fine-tuning a base model on an instruction or chat dataset with SFTTrainer before further alignment.
  • Aligning a model to human or synthetic preference data with DPOTrainer or KTOTrainer instead of running full RLHF with a separate reward model.
  • Training reasoning models with GRPOTrainer against verifiable reward functions, the method used to train models like DeepSeek-R1.
  • Training a reward model with RewardTrainer to score generations for a downstream RLHF or best-of-n pipeline.
  • Distilling a large teacher model into a smaller student model via DistillationTrainer.

Under The Hood

Architecture Every trainer (SFTTrainer, DPOTrainer, GRPOTrainer, KTOTrainer, RewardTrainer, RLOOTrainer, DistillationTrainer) subclasses a shared _BaseTrainer that itself wraps transformers.Trainer, layering method-specific loss computation, data collation, and generation logic on top of Transformers’ existing training loop rather than reimplementing it; each trainer pairs with a matching Config dataclass (e.g. GRPOConfig, DPOConfig) that extends transformers.TrainingArguments, keeping the hyperparameter surface consistent across methods. The trl.experimental namespace mirrors this pattern for unstable methods (async GRPO, distillation, ORPO, and more) using the same base-trainer/config contract, so new post-training algorithms can be added without touching the stable trainer classes, and adding a new method means writing one trainer/config pair rather than modifying shared core logic.

Tech Stack TRL is pure Python (3.10–3.14) built with setuptools, depending on Hugging Face’s own stack — Transformers for models/tokenizers, Accelerate for distributed execution, Datasets for data loading, and PyTorch as the underlying tensor/autograd engine. Optional extras wire in PEFT for LoRA/QLoRA, bitsandbytes for quantization, a version-pinned vLLM plus supporting HTTP libraries for fast generation during GRPO-style rollouts, DeepSpeed, Liger kernels for fused ops, and specialized reward-scoring extras. A console-script entry point exposes the trl CLI, and CI runs the suite across multiple Python versions and against Transformers’ own development branch to catch upstream breakage early.

Code Quality The test suite is extensive, with dozens of files covering each trainer and loss function individually plus dedicated distributed, experimental, and invariant-testing subdirectories, run via pytest with coverage and flake-tolerant reruns, and a separate CI workflow gates the experimental namespace. Core modules use modern typed Python with dataclass-based configs, clear docstrings, and consistent license headers; pre-commit hooks and a documentation-build extra enforce formatting and doc quality before merge. Error handling favors explicit warnings and guard assertions over silent fallbacks in the trainer code reviewed.

What Makes It Unique TRL’s distinguishing choice is shipping new post-training research — GRPO, KTO, RLOO, and several experimental variants — as production-ready trainer subclasses within a short window of publication, with the experimental namespace formalizing an explicit path from research code to stable API. Its GRPOTrainer integrates a separate inference engine directly into the training loop for fast on-policy generation, a nontrivial engineering problem that most from-scratch RLHF implementations avoid solving.

Used by 6 apps in this directory

Python
59%
Apache 2.0

argilla

AI Development · Data Engineering

5,125

Collaborate on high-quality AI training data with a self-hosted annotation platform built for LLMs, NLP, and multimodal models.

View details
65
Repo Health
81
Technical
61
Dependency
Built with
Python 59%
Jupyter Notebook 21%
Updated 1 weeks ago
Python
92%
Apache 2.0

ART

AI Development

10,779

Give your LLM agents on-the-job training—ART lets you apply GRPO reinforcement learning to any multi-step agentic workflow with minimal code changes.

View details
85
Repo Health
82
Technical
73
Dependency
Built with
Python 92%
Updated 6 days ago
Python
62%
Apache 2.0

marimo

Data Engineering · Developer Tools

22,918

A reactive Python notebook that eliminates hidden state, runs reproducibly, and deploys as a web app or script — stored as pure Python, built for the AI era.

View details
89
Repo Health
91
Technical
65
Dependency
Built with
Python 62%
TypeScript 37%
Updated 6 days ago
Rust
63%
MIT

PostgresML

AI Development · Databases

6,825

Run ML training and LLM inference natively inside PostgreSQL with GPU acceleration — no data movement required.

View details
51
Repo Health
76
Technical
64
Dependency
Built with
Rust 63%
JavaScript 11%
Updated 1 years ago
Python
76%
Apache 2.0

Second Me

AI Assistants · Productivity

15,690

Train a locally hosted AI twin on your own memories—then connect it to the world through a decentralized identity network.

View details
41
Repo Health
75
Technical
67
Dependency
Built with
Python 76%
TypeScript 19%
Updated 1 years ago
Python
73%
Apache 2.0

Unsloth

AI Assistants · AI Development

76,886

Run and fine-tune LLMs, diffusion, audio and embedding models on your own hardware, from a native desktop app, a browser UI, or a Python library.

View details
89
Repo Health
83
Technical
70
Dependency
Built with
Python 73%
TypeScript 21%
Updated 5 days ago

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers