← back to the directory
Library / SDK Inference & ServingFine-tuning & Training

Optimum

1Optimizes 2Transformers 3models 4for 5target 6hardware

the six Ws · specification

W1 Who

ML engineers deploying Hugging Face models who need faster or smaller inference on specific hardware.

W2 What

A Hugging Face toolkit that exports, quantizes, and accelerates Transformers, Diffusers, and Sentence Transformers models for ONNX Runtime, OpenVINO, TensorRT, and specialized chips.

W3 Where

Used in Python pipelines targeting CPUs, GPUs, Intel, AWS Inferentia, and other accelerators.

W4 When

Use when a trained model needs to be compressed or exported before production deployment.

W5 Why

It closes the gap between a research checkpoint and an efficient production inference target.

W6 With

Built on top of Hugging Face Transformers and integrates with ONNX Runtime, OpenVINO, and vendor SDKs.

W7 Watch

Open source, Apache 2.0 licensed, maintained by Hugging Face; no credentials required for local use.

quantizationonnx exporthardware accelerationopenvinotensorrt

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.