← back to the directory
Model Inference & Serving

Pixtral 12B

1Mistral's 2first 3open 4multimodal 5vision 6model

the six Ws · specification

W1 Who

Mistral AI released Pixtral 12B, its first vision-language model.

W2 What

A 12-billion-parameter multimodal model that interleaves images and text for natural-image reasoning and chat.

W3 Where

Distributed on Hugging Face and via Mistral's own model hub.

W4 When

Pixtral 12B was released in September 2024.

W5 Why

It extended Mistral's open-weight lineup into multimodal understanding, competing with Qwen2-VL and LLaVA.

W6 With

Runs via the Mistral inference stack or vLLM/transformers, using a new vision encoder trained from scratch.

W7 Watch

Open-weight release by Mistral AI under Apache 2.0 license.

vision-languagemultimodalmistralopen-weights

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.