We put the same work to every large language model. Then we ranked them.
One identical set of tasks, 53 models, scored on craft, speed, control and value. Closed APIs and open weights on the same page, because that is the choice.
- Tracked
- 53
- Fully reviewed
- 00
- Categories
- 05
- Average score
- 79.1
The short answer
What is the best AI model in 2026?
Claude Fable 5 leads our index of 53 large language models with 89.2 out of 100, ahead of GPT-5.6 Sol on 88.6 and Claude Opus 5 on 88.5. Models are scored on answer quality, latency, steerability and capability per dollar, with answer quality weighted heaviest. Nothing here is sponsored.
Index updated
The model index
Every large language model we track, in rank order.
Frontier flagships, open weights you can host yourself, and small models quick enough to sit in a loop. Narrow by what kind of model it is, or re-sort by what you care about.
53 builders
0189.2Claude Fable 5
The best answer money can buy, and it is a lot of money.
FrontierProvisional
0288.6GPT-5.6 Sol
OpenAI's deepest thinker: the one to hand a problem you cannot solve yourself.
ReasoningProvisional
0388.5Claude Opus 5
Most of the flagship, at half the price. The one to reach for by default.
FrontierProvisional
0487.9Claude Sonnet 5
The coding workhorse: near the top on real software tasks, at a working price.
CodingProvisional
0587.5Kimi K3
The largest open-weights model anyone has shipped, and it reads a million tokens without blinking.
Open weightsProvisional
- GLM-5.20687.1
GLM-5.2
Frontier-class output, an MIT licence, and a million tokens of context. That combination is the story.
Open weightsProvisional
- DeepSeek V4 Pro0786.8
DeepSeek V4 Pro
A reasoning model at the frontier, published under MIT, priced like an afterthought.
ReasoningProvisional
- Gemini 3.6 Flash0886.6
Gemini 3.6 Flash
Frontier answers at Flash speed, with a free tier you can actually build on.
FrontierProvisional
- Grok 4.50985.0
Grok 4.5
Strong, current, and unusually willing to answer the question you actually asked.
FrontierProvisional
- MiniMax M31084.5
MiniMax M3
Open weights tuned for software work, and the fastest decoding of anything in its class.
CodingProvisional
1184.4GPT-5.6 Terra
The middle of OpenAI's range, and the one most teams should actually deploy.
FrontierProvisional
- Qwen 3.61283.3
Qwen 3.6
The open Qwen release most people should actually run, in sizes that fit real hardware.
Open weightsProvisional
1383.3GPT-5.6 Luna
Frontier-family manners at a price that lets you stop counting tokens.
Small and fastProvisional
1482.9Mistral Large 3
Europe's answer, published under Apache 2.0, with a quarter of a million tokens of context.
Open weightsProvisional
- DeepSeek V4 Flash1582.6
DeepSeek V4 Flash
Near-frontier answers at roughly a hundredth of frontier prices, with the weights published.
Open weightsProvisional
The complete index
Every entry on this page, in rank order. The cards above load in batches; this list does not.
- 01Claude Fable 589.2
- 02GPT-5.6 Sol88.6
- 03Claude Opus 588.5
- 04Claude Sonnet 587.9
- 05Kimi K387.5
- 06GLM-5.287.1
- 07DeepSeek V4 Pro86.8
- 08Gemini 3.6 Flash86.6
- 09Grok 4.585.0
- 10MiniMax M384.5
- 11GPT-5.6 Terra84.4
- 12Qwen 3.683.3
- 13GPT-5.6 Luna83.3
- 14Mistral Large 382.9
- 15DeepSeek V4 Flash82.6
- 16Xiaomi MiMo V2.5 Pro82.4
- 17Qwen 3.5 397B81.9
- 18Step 3.7 Flash81.6
- 19Doubao Seed 2.0 Pro81.5
- 20Muse Spark81.3
- 21Qwen 3.8-Max81.2
- 22Qwen 3.7-Max81.0
- 23Llama 4 Maverick80.4
- 24Nemotron 3 Super80.1
- 25Gemma 480.1
- 26Hunyuan HY380.0
- 27Solar Open 279.2
- 28Claude Haiku 4.578.9
- 29Mistral Medium 3.578.6
- 30Llama 4 Behemoth78.2
- 31Mistral Small 478.1
- 32EXAONE 2.077.8
- 33ERNIE 5.177.7
- 34openPangu 2.0 Pro77.7
- 35Llama 4 Scout77.3
- 36Amazon Nova Premier77.3
- 37Hermes 477.0
- 38Grok 4.2076.9
- 39Cohere Command A+76.8
- 40Phi-4 Reasoning Vision76.8
- 41Gemini 3.5 Flash-Lite75.2
- 42Sarvam 105B74.3
- 43Mercury 273.2
- 44IBM Granite 3.272.5
- 45Amazon Nova Lite72.4
- 46OLMo 3.172.0
- 47Perplexity Sonar Pro71.6
- 48Doubao Seed 1.6 Flash71.2
- 49HyperCLOVA X THINK70.7
- 50Ministral 369.0
- 51Reka Edge68.7
- 52SmolLM367.7
- 53Liquid LFM 266.7
Weekly update
The ranking moves. Find out when.
One short email a week: what changed in the index, which builder moved and why, and anything new worth trying. No sponsorships, no reprinted press releases.