← back to the directory
Library / SDK Inference & Serving

KTransformers

1Runs 2giant 3MoE 4models 5on 6GPU

the six Ws · specification

W1 Who

Researchers and engineers running huge mixture-of-experts models on limited local GPU hardware.

W2 What

KTransformers is an open-source hybrid CPU-GPU inference and fine-tuning framework that offloads MoE expert weights to RAM so models like DeepSeek-671B can run on a single consumer GPU.

W3 Where

Self-hosted, run locally from source or pip on Linux workstations with a GPU and large system RAM.

W4 When

Released by Tsinghua's MADSys Lab, with SOSP 2025 publication and active development into 2026.

W5 Why

It exists to make trillion-parameter-class MoE inference feasible without datacenter-scale GPU memory.

W6 With

Built on PyTorch and integrates optimized kernels from llama.cpp and Marlin.

W7 Watch

Open source, Apache-2.0 license, self-hosted and runs fully local; no API key or external data transfer required.

CPU-GPU offloadingMoE inferenceheterogeneous computingsingle-GPU deployment

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.