← back to the directory
Library / SDK Inference & ServingCoding

Candle

1Minimalist 2Rust 3framework 4for 5ML 6inference

the six Ws · specification

W1 Who

Rust developers wanting to run or deploy ML models without a Python runtime dependency.

W2 What

Candle is a Hugging Face maintained minimalist machine learning framework for Rust with CPU and CUDA GPU backends and WebAssembly support.

W3 Where

Used as a Rust crate embedded directly in applications and servers, or compiled to WebAssembly for browser deployment.

W4 When

Released by Hugging Face around 2023 and actively maintained as of 2026.

W5 Why

Enables fast, memory safe, dependency light model inference for LLMs and diffusion models in performance sensitive or embedded Rust environments.

W6 With

Distributed as a Rust crate via crates.io, with example model implementations for LLaMA, Whisper, and Stable Diffusion in the same repository.

W7 Watch

Open source under a dual Apache-2.0 and MIT license, maintained by Hugging Face; runs fully locally or in browser with no required external API calls.

RustML frameworkGPU inferenceWASMmodel inference

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.