DeepSeek V4 Flash
DeepSeek V4 Flash
Near-frontier answers at roughly a hundredth of frontier prices, with the weights published.
Updated
The read
Flash is the cheap half of DeepSeek V4 and the reason the whole market repriced. It is a sparse model that activates a small slice of its parameters per token, which is why it is both quick and almost free to call, and the weights are MIT so you are not locked to the API at all. It gives up depth on the hardest problems. On everything else the price is difficult to argue with.
Where it shines
- Among the cheapest capable models in the world per token
- MIT weights, so the hosted API is a convenience rather than a lock-in
- Fast, thanks to a small active parameter count
Worth knowing first
- Noticeably shallower than the Pro tier on hard reasoning
- Some organisations will not send data to a Chinese-hosted API
Index score
82.6
Provisional. This score is read from the product's public capability, not from the one-prompt rebuild the fully reviewed entries went through, so treat it as a placing rather than a verdict.
- Rank
- 15 of 53
- Builds
- LLMs
- Category
- Open weights
- Output
- MIT-licensed weights, plus a very cheap API
- Pricing
- Free tier
- Best for
- High-volume work where the token bill is the constraint
Scores are our own editorial judgement, weighted craft 55, speed 10, control 15 and value 20.
Also worth a look
Full index- 05Kimi K3Kimi K3ProvisionalThe largest open-weights model anyone has shipped, and it reads a million tokens without blinking.Open weightsFree tier87.5
- 06GLM-5.2GLM-5.2ProvisionalFrontier-class output, an MIT licence, and a million tokens of context. That combination is the story.Open weightsFree tier87.1
- 12Qwen 3.6Qwen 3.6ProvisionalThe open Qwen release most people should actually run, in sizes that fit real hardware.Open weightsFree tier83.3