Pixtral 12B
the six Ws · specification
Mistral AI released Pixtral 12B, its first vision-language model.
A 12-billion-parameter multimodal model that interleaves images and text for natural-image reasoning and chat.
Distributed on Hugging Face and via Mistral's own model hub.
Pixtral 12B was released in September 2024.
It extended Mistral's open-weight lineup into multimodal understanding, competing with Qwen2-VL and LLaVA.
Runs via the Mistral inference stack or vLLM/transformers, using a new vision encoder trained from scratch.
Open-weight release by Mistral AI under Apache 2.0 license.