← back to the directory
Library / SDK Fine-tuning & TrainingInference & Serving

DeepSpeed

1Scale 2and 3speed 4up 5model 6training.

the six Ws · specification

W1 Who

Built by Microsoft.

W2 What

Provides ZeRO optimizer sharding and offload for large-scale training and inference.

W3 Where

A Python library on multi-GPU systems.

W4 When

When training exceeds a single GPU's memory.

W5 Why

Trains models too large for one device.

W6 With

PyTorch and CUDA.

W7 Watch

Microsoft open-source, Apache-2.0; runs on your GPUs, no credentials.

MicrosoftZeROParallelism

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.