← back to the directory
Library / SDK Fine-tuning & Training

nanotron

1Hugging 2Face's 3minimalistic 43D-parallelism 5pretraining 6library

the six Ws · specification

W1 Who

Hugging Face develops nanotron as a minimalistic large-scale training library.

W2 What

nanotron implements 3D parallelism (tensor, pipeline, and data) for pretraining transformer language models from scratch.

W3 Where

Hosted on GitHub under the huggingface organization.

W4 When

First released publicly in 2024 and under active development.

W5 Why

Built to give researchers a simple, hackable alternative to heavier distributed-training frameworks.

W6 With

Runs via torchrun with YAML configuration files across multi-GPU clusters, built on PyTorch.

W7 Watch

Open source under Apache 2.0, self-hosted, no credentials required; maintained by Hugging Face.

3D-parallelismdistributed-trainingpretrainingPyTorchlarge-scale-LLM

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.