LLMs 101

Monthly model tracker

How AI models rank when
real humans compare them

Benchmark scores don't tell the full story. This tracker translates human preference data and real-world usage into plain-English tiers — updated every month.

Updated 16th August 2026
🏆
What is human preference ranking?
Real users are shown two anonymous AI responses to the same prompt and pick the one they prefer. Thousands of these battles produce an Elo-style ranking — the same system used to rank chess players.
💰
What does "cost vibe" mean?
Instead of listing raw dollar-per-token figures (which change constantly), we use plain-English cost tiers: Free, Low, Standard, and Premium — reflecting how expensive it is to run each model at real production scale.
📊
Why not show benchmark scores?
Benchmarks like MMLU or HumanEval measure narrow academic tasks. A score of "86.2%" means nothing in practice. Human preference rankings and real use-case guidance are far more useful for most readers.
🔄
How often is this updated?
Monthly. The AI landscape moves fast — a model released last month can jump several tiers. Check back each month, or read the Trends section for major shifts as they happen.
Show
How to read this table: Rankings reflect our monthly editorial review of current AI benchmarks and independent testing, combined with qualitative assessment of real-world professional use. Tier 1 = consistently preferred across that review. Tiers are relative — all listed models are genuinely excellent by any historical standard.
1
Anthropic
Most capable general-access model
🏆 Tier 1 - Top
Best for
State-of-the-art performance on the hardest, longest-running coding, research, and knowledge-work tasks — software engineering, scientific/financial analysis, and multi-day autonomous agent work where raw capability matters more than cost.
Premium Closed API
2
Anthropic
Flagship intelligence model
🥈 Tier 1 - Top
Best for
Serious coding, long-running agentic work, and high-stakes enterprise and legal tasks where reliability and reasoning quality matter most.
Premium Closed API
3
OpenAI
Flagship reasoning and coding model
🥉 Tier 1 - Top
Best for
Complex professional work across coding, science, and research. OpenAI's most capable model and the strongest reasoning option in ChatGPT.
Premium Closed API
4
Google DeepMind
Advanced multimodal reasoning
Tier 1 - Top
Best for
Deep reasoning over huge multimodal inputs. Its 1M-token context handles entire codebases, long documents, audio, and video in a single prompt.
Premium Closed API
5
xAI
Efficient coding and agentic model
Tier 1 - Efficient frontier
Best for
Real software engineering and agentic tasks. Trained alongside Cursor, it is pitched as Opus-class quality at faster speeds and lower cost.
Standard Closed API
6
Anthropic
Balanced everyday Claude
Tier 2 - Best value Claude
Best for
A lower-cost Claude for agentic coding, tool use, and knowledge work when you want strong quality without paying Opus-tier prices.
Standard Closed API
7
OpenAI
Balanced value model
Tier 2 - Best value GPT
Best for
Everyday work that needs solid intelligence at lower cost. Terra matches the prior GPT-5.5 flagship on benchmarks at roughly half the price.
Standard Closed API
8
Z.ai
Top open-weight model
Tier 3 - Open-weight leader
Best for
Self-hosted or low-cost agentic coding. The highest-ranked open-weight model on the Artificial Analysis Intelligence Index, MIT-licensed with 1M context.
Low Open weights
9
Moonshot AI
Largest open-weight model
Tier 3 - Open-weight frontier
Best for
Long-horizon coding and agent workflows. At 2.8T parameters with a 1M-token context, it is the largest openly downloadable model to date.
Low Open weights
10
DeepSeek
Cost-efficient open frontier
Tier 3 - Open-weight value
Best for
Frontier-adjacent chat, coding, and agent work at very low cost. A 1.6T mixture-of-experts model with a 1M-token context, self-hostable via open weights.
Ultra-low Open weights
11
Google DeepMind
Fast, cheap workhorse
Tier 4 - Speed and cost leader
Best for
High-volume coding, knowledge work, and multimodal tasks where speed and price matter. Google's current default Flash workhorse, cheaper than its predecessor.
Ultra-low Closed API
12
OpenAI
Fastest, lowest-cost GPT
Tier 4 - Budget and high-volume
Best for
Cost-sensitive, high-throughput workloads like classification, routing, and simple tool calls. The fastest and cheapest model in the current GPT-5.6 family.
Ultra-low Closed API
13
Alibaba Qwen
Open coding specialist
Tier 3 - Open coding specialist
Best for
Local, self-hosted coding. The dense 27B model beats far larger predecessors on agentic coding benchmarks and runs on a single consumer GPU under Apache 2.0.
Ultra-low Open weights

* Llama models are free to download and run but require your own hardware or cloud compute. Self-hosting costs (electricity, GPU rental) vary. · Rankings reflect monthly editorial review of current AI benchmarks and independent testing, combined with editorial assessment — see our Methodology & Glossary page for full detail. · Cost tiers are illustrative and based on that same monthly review — actual pricing changes frequently; check each provider's current pricing page for exact figures. · This tracker focuses on general-purpose chat and reasoning models. Specialised models (image generation, audio, video) are not included.