Llama 4 Scout
Llama 4 Scout
Small, multimodal and open, for when the model has to live on your own box.
Updated
The read
Scout is the small end of Llama 4: a mixture-of-experts model light enough to serve on a single high-memory GPU while still reading images. For an internal tool, an on-premise assistant, or anything where the model has to sit next to the data, it is the least complicated open option. It is a small model, and it answers like one when the question gets hard.
Where it shines
- Serves on a single high-memory GPU
- Multimodal despite its size
- Drop-in support in every open serving stack
Worth knowing first
- Reasoning is thin next to the larger models in the same family
- The community licence carries a user-count threshold
Index score
77.3
Provisional. This score is read from the product's public capability, not from the one-prompt rebuild the fully reviewed entries went through, so treat it as a placing rather than a verdict.
- Rank
- 35 of 53
- Builds
- LLMs
- Category
- Open weights
- Output
- Downloadable weights under Meta's community licence
- Pricing
- Free tier
- Best for
- On-premise assistants that must run on one GPU
Scores are our own editorial judgement, weighted craft 55, speed 10, control 15 and value 20.
Also worth a look
Full index- 05Kimi K3Kimi K3ProvisionalThe largest open-weights model anyone has shipped, and it reads a million tokens without blinking.Open weightsFree tier87.5
- 06GLM-5.2GLM-5.2ProvisionalFrontier-class output, an MIT licence, and a million tokens of context. That combination is the story.Open weightsFree tier87.1
- 12Qwen 3.6Qwen 3.6ProvisionalThe open Qwen release most people should actually run, in sizes that fit real hardware.Open weightsFree tier83.3