← back to the directory
Library / SDK Inference & Serving

IPEX-LLM

1Accelerates 2LLMs 3on 4Intel 5CPUs 6GPUs

the six Ws · specification

W1 Who

Developers running LLMs locally on Intel CPUs, integrated GPUs, NPUs, or Arc/Flex/Max discrete GPUs.

W2 What

IPEX-LLM is a PyTorch library that accelerates local LLM inference and fine-tuning on Intel hardware using low-bit quantization and XPU-specific optimizations.

W3 Where

Self-hosted, installed as a Python package on Intel-equipped PCs and servers.

W4 When

Formerly BigDL-LLM, actively maintained by Intel with continuous updates through 2026.

W5 Why

It exists to make Intel client and server chips competitive for local LLM workloads without dedicated NVIDIA GPUs.

W6 With

Integrates with llama.cpp, Ollama, Hugging Face Transformers, LangChain, and vLLM.

W7 Watch

Open source, Apache-2.0 license, self-hosted; no API key needed; all inference runs on local Intel hardware.

Intel XPUlow-bit quantizationCPU/iGPU inferencePyTorch integration

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.