Skip to content

Shortlist

What is the cheapest capable LLM in 2026?

DeepSeek V4 Flash is our pick, scoring 82.6 out of 100 across craft, speed, control and value, with Gemma 4 behind it on 80.1. These are the five best free-tier models that are still strong on answer quality, chosen from 19 that qualify and ranked on the same rubric as the rest of the index.

Index updated

Chosen from 19 that qualify, out of 53 tracked

The shortlist

  1. 01

    DeepSeek V4 Flash

    Near-frontier answers at roughly a hundredth of frontier prices, with the weights published.

    Open weightsFree tier

    Best for

    High-volume work where the token bill is the constraint

    82.6

  2. 02

    Gemma 4

    Google's open weights, from something that runs on a phone up to a real server model.

    Open weightsFree tier

    Best for

    Running a capable model on hardware you control

    80.1

  3. 03

    GLM-5.2

    Frontier-class output, an MIT licence, and a million tokens of context. That combination is the story.

    Open weightsFree tier

    Best for

    Frontier-class work you need to keep in house

    87.1

  4. 04

    MiniMax M3

    Open weights tuned for software work, and the fastest decoding of anything in its class.

    CodingFree tier

    Best for

    Coding agents that need speed as much as accuracy

    84.5

  5. 05

    Qwen 3.6

    The open Qwen release most people should actually run, in sizes that fit real hardware.

    Open weightsFree tier

    Best for

    Self-hosting and fine-tuning on hardware you already have

    83.3

More shortlists on this index