← back to the directory
Library / SDK Inference & Serving

Nexa SDK

1On-device 2inference 3across 4CPU 5GPU 6NPU

the six Ws · specification

W1 Who

App developers deploying LLMs and multimodal models directly on phones, PCs, and edge devices.

W2 What

Nexa SDK is an on-device inference framework running LLMs, VLMs, ASR, and TTS models across CPU, GPU, and NPU backends with an OpenAI-compatible local API server.

W3 Where

Self-hosted on Android, iOS, Windows, macOS, and Linux/IoT devices via native bindings.

W4 When

Actively maintained by Nexa AI with day-0 support for new model releases through 2026.

W5 Why

It exists to bring efficient, private, quantized model inference to consumer and edge hardware.

W6 With

Supports GGUF, MLX, and Nexa's own quantized model formats.

W7 Watch

Open source SDK with commercial hardware partnerships; self-hosted, runs fully on-device; no data leaves the device.

on-device inferenceNPU accelerationcross-platformGGUF MLX supportmultimodal

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.