Library / SDK
Fine-tuning & TrainingInference & Serving
DeepSpeed
1Scale 2and 3speed 4up 5model 6training.
the six Ws · specification
W1
Who
Built by Microsoft.
W2
What
Provides ZeRO optimizer sharding and offload for large-scale training and inference.
W3
Where
A Python library on multi-GPU systems.
W4
When
When training exceeds a single GPU's memory.
W5
Why
Trains models too large for one device.
W6
With
PyTorch and CUDA.
W7
Watch
Microsoft open-source, Apache-2.0; runs on your GPUs, no credentials.
alternatives
works with