Shortlist
What is the cheapest capable LLM in 2026?
DeepSeek V4 Flash is our pick, scoring 82.6 out of 100 across craft, speed, control and value, with Gemma 4 behind it on 80.1. These are the five best free-tier models that are still strong on answer quality, chosen from 19 that qualify and ranked on the same rubric as the rest of the index.
Index updated
Chosen from 19 that qualify, out of 53 tracked
The shortlist
- 01
DeepSeek V4 Flash
Near-frontier answers at roughly a hundredth of frontier prices, with the weights published.
Open weightsFree tierBest for
High-volume work where the token bill is the constraint
- 02
Gemma 4
Google's open weights, from something that runs on a phone up to a real server model.
Open weightsFree tierBest for
Running a capable model on hardware you control
- 03
GLM-5.2
Frontier-class output, an MIT licence, and a million tokens of context. That combination is the story.
Open weightsFree tierBest for
Frontier-class work you need to keep in house
- 04
MiniMax M3
Open weights tuned for software work, and the fastest decoding of anything in its class.
CodingFree tierBest for
Coding agents that need speed as much as accuracy
- 05
Qwen 3.6
The open Qwen release most people should actually run, in sizes that fit real hardware.
Open weightsFree tierBest for
Self-hosting and fine-tuning on hardware you already have