Mixtral 8x7B
the six Ws · specification
Mistral AI released Mixtral 8x7B, its first sparse mixture-of-experts model.
A 47-billion-parameter (12.9B active) MoE language model matching or beating Llama 2 70B and GPT-3.5 on many benchmarks.
Distributed on Hugging Face and via Mistral's API.
Mixtral 8x7B was released in December 2023, with Mixtral 8x22B following in April 2024.
It proved sparse MoE architectures could deliver top-tier performance at lower inference cost, popularizing the approach in open models.
Runs via vLLM, transformers, or llama.cpp, requiring MoE-aware inference routing.
Open-weight release by Mistral AI under Apache 2.0 license.
alternatives
works with