Mistral Large 3
Mistral Large 3
Europe's answer, published under Apache 2.0, with a quarter of a million tokens of context.
Updated

The read
Large 3 is a sparse model that activates a fraction of its parameters per token, which is how Mistral gets frontier-adjacent quality out of something you can actually host. The licence is Apache 2.0 with no conditions attached, the context runs to a quarter of a million tokens, and for European teams it answers the data-residency question without a legal review. It is a step behind the very top on hard reasoning.
Where it shines
- Apache 2.0 with no user, revenue or field-of-use clauses
- Sparse design keeps serving costs sane for its quality
- EU-hosted API option, which several sectors require
Worth knowing first
- Behind the leading open models from Chinese labs on benchmarks
- Latency is unremarkable for a model of its size
Index score
82.9
Provisional. This score is read from the product's public capability, not from the one-prompt rebuild the fully reviewed entries went through, so treat it as a placing rather than a verdict.
- Rank
- 14 of 53
- Builds
- LLMs
- Category
- Open weights
- Output
- Apache-licensed weights, plus a hosted API
- Pricing
- Free tier
- Best for
- European teams that need capability and data residency at once
Scores are our own editorial judgement, weighted craft 55, speed 10, control 15 and value 20.
Compare it with
Also worth a look
Full index- 05Kimi K3Kimi K3ProvisionalThe largest open-weights model anyone has shipped, and it reads a million tokens without blinking.Open weightsFree tier87.5
- 06GLM-5.2GLM-5.2ProvisionalFrontier-class output, an MIT licence, and a million tokens of context. That combination is the story.Open weightsFree tier87.1
- 12Qwen 3.6Qwen 3.6ProvisionalThe open Qwen release most people should actually run, in sizes that fit real hardware.Open weightsFree tier83.3