← back to the directory
Library / SDK Fine-tuning & Training

TRL

1Train 2language 3models 4with 5reinforcement 6learning.

the six Ws · specification

W1 Who

Maintained by Hugging Face.

W2 What

Provides SFT, DPO, PPO, and GRPO trainers for language models.

W3 Where

A Python library running on your GPUs.

W4 When

When post-training or aligning an open model.

W5 Why

The standard toolkit for RLHF and preference tuning.

W6 With

PyTorch, Transformers, PEFT, and Accelerate.

W7 Watch

Official Hugging Face library, Apache-2.0, no credentials; trains on your own hardware.

RLHFDPOPPOHugging-Face

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.