← back to the directory
Library / SDK Fine-tuning & Training

OpenRLHF

1A 2scalable 3RLHF 4framework 5on 6Ray.

the six Ws · specification

W1 Who

By the OpenRLHF community.

W2 What

A high-performance RLHF framework combining Ray, vLLM, and DeepSpeed.

W3 Where

A Python framework on GPU clusters.

W4 When

When doing RLHF at real scale.

W5 Why

Production-ready, distributed preference training.

W6 With

Python, Ray, vLLM, and DeepSpeed.

W7 Watch

Apache-2.0 open-source; runs on your cluster, no credentials.

RLHFRayvLLM

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.