LLMs 101

Frontier model directory

Every major AI model,
actually explained

No benchmark scores. No jargon. Just what each model family is genuinely good at, where it falls short, and who should use it.

Updated 22nd July 2026
Filter by
OpenAI Market Leader
GPT-5.6 Sol (flagship) · GPT-5.6 Terra · GPT-5.6 Luna · GPT-5.5 (Thinking / Pro) · GPT-5.5 Instant · GPT-5.4
Core superpower
A three-tier flagship family that lets you dial cost against capability — from top-end reasoning to fast, cheap, high-volume work — without switching brands.
Key trade-off
The tiers can be genuinely confusing: "available in ChatGPT" doesn't mean every model shows up in your picker, and access depends on your plan, effort settings, and rollout wave.
Speed
8/10
Reasoning
9.5/10
Cost
Premium
Best non-technical use
Working through a hard, multi-step task — drafting a detailed report, planning a project, or reasoning through a tricky decision — where you want the strongest thinking available in ChatGPT.
Cost tier
Premium

OpenAI's current flagship family is GPT-5.6, which was released publicly on July 9, 2026 across ChatGPT, the API, and Codex as three models: Sol, the frontier flagship for complex reasoning, coding, and long-horizon agentic work; Terra, the balanced everyday model; and Luna, the fastest and cheapest tier built for high-volume workloads. It followed a limited preview that began on June 26, 2026 for a small group of trusted partners, with access confined to the API, Codex, or both. All three models share the same core specs: a 1 million token context window, 128,000 maximum output tokens, and a February 2026 knowledge cutoff. In practice, eligible paid plans access GPT-5.6 Sol through ChatGPT's reasoning settings while GPT-5.5 Instant remains the default for everyday chat, and Terra and Luna are available in Work, Codex, and the API depending on the product and plan. API pricing runs from $5 input / $30 output per million tokens for Sol, $2.50 / $15 for Terra, and $1 / $6 for Luna.

Anthropic Market Leader
Fable 5, Opus 4.8, Sonnet 5, Haiku 4.5
Core superpower
Careful, nuanced reasoning and long-form writing that stays coherent across very long, complex tasks.
Key trade-off
Top-tier models sit at a premium price, and Fable 5 access briefly hinged on shifting US export rules — a reminder that frontier availability can change fast.
Speed
7/10
Reasoning
9.6/10
Cost
Premium
Best non-technical use
Drafting, editing, and thinking through long documents — reports, proposals, research summaries — where tone and accuracy matter.
Cost tier
Premium

Claude is Anthropic's family of AI models, spanning the flagship Fable 5, the Opus 4.8 reasoning tier, and the faster Sonnet and Haiku variants. As of July 21, 2026, Claude Fable 5 is available globally after US export controls imposed on June 12 were lifted on June 30 and access was restored on July 1; Anthropic describes it as its most capable widely released model, sitting above the Opus tier. Claude is widely used for complex reasoning, coding, and long-form writing, with a context window of up to 1 million tokens on its Opus 4.8 and Sonnet 5 models at flat per-token pricing (Haiku 4.5 supports 200,000 tokens).

Google DeepMind Highly Competitive
Gemini 3.5 Flash, Gemini 3.1 Pro, Gemini 3.1 Flash-Lite
Core superpower
Gemini 3.5 Flash packs near-Pro coding and agentic ability into a fast, high-throughput package, so it can churn through multi-step tool-use tasks quickly and cheaply.
Key trade-off
It is fast for its intelligence class but not the outright speed leader — and on deep reasoning, long-context retrieval, and knowledge-heavy work, the pricier Gemini 3.1 Pro (and rivals from OpenAI and Anthropic) still pull ahead.
Speed
8.5/10
Reasoning
8.2/10
Cost
Good value
Best non-technical use
Running an AI assistant that works through long, multi-step tasks — like reading a stack of invoices or a 100-page document and pulling out the answers — where you want speedy, affordable results and can accept slightly-below-frontier accuracy.
Cost tier
Low

Gemini is Google DeepMind's family of multimodal AI models, spanning the fast, efficient Gemini 3.5 Flash, the higher-intelligence Gemini 3.1 Pro, and the ultra-cheap Gemini 3.1 Flash-Lite. Gemini 3.5 Flash offers a 1 million–token context window and is optimized for high-throughput coding and agentic ("tool-using") workflows, scoring 55 on the Artificial Analysis Intelligence Index while running faster than most models in its intelligence tier. It is a strong pick for high-volume, multi-step tasks like document processing, research assistants, and automated workflows where speed and cost matter as much as raw capability.

xAI Highly Competitive
Grok 4.5 · Grok 4.3 · Grok 4.1 Fast
Core superpower
Live access to what people are posting on X right now, so it's unusually good at "what's happening lately" questions that stump models working from a fixed training set.
Key trade-off
The newest model has the smallest memory: Grok 4.5 reads about 500K tokens at once, while the older Grok 4.3 handles 1M and Grok 4.1 Fast handles 2M — so for very long documents, the newest isn't always the right pick.
Speed
8/10
Reasoning
8/10
Cost
Moderate
Best non-technical use
Getting a fast, current read on breaking news or public reaction — asking what people are saying about a company, event, or trend as it unfolds.
Cost tier
Standard

Grok is xAI's family of models, with Grok 4.5 as the current flagship since its July 8, 2026 release — a coding- and agent-focused model with a 500K-token context window, priced at $2 per million input tokens and $6 per million output. Below it sit Grok 4.3, the lower-cost 1M-context tier launched in April 2026, and Grok 4.1 Fast, a budget, high-throughput option with a 2M-token window. Grok's defining feature is real-time access to X, making it a strong choice for questions that depend on the latest public conversation rather than a static training cutoff.

Meta AI Rising
Llama 4 Maverick · Llama 4 Scout · Llama 3.3 70B
Core superpower
Fully open weights — download and run privately on your own hardware, free forever with no API costs; Llama 4 Scout's 10-million-token context window is the largest of any openly available model
Key trade-off
Meta has shifted its frontier investment to a new proprietary model (Muse Spark, below) — Llama remains available but is expected to see maintenance updates rather than continued frontier development, and now trails closed frontier models by a wider margin than before
Speed
7.3/10
Reasoning
4.8/10
Cost
Free*
Best non-technical use
Privacy-sensitive workflows where data cannot leave your machine, high-volume automation where API costs would otherwise be prohibitive
Cost tier
Free

Meta's Llama 4 family remains the most widely deployed open-weight AI ecosystem. Llama 4 Maverick is a natively multimodal Mixture-of-Experts model with a 1-million-token context window. Llama 4 Scout offers the same multimodal capability with a 10-million-token context window — the largest of any openly available model. Llama 3.3 70B remains a widely-deployed practical text-only option. In April 2026, Meta Superintelligence Labs launched a new proprietary model, Muse Spark, as the engine behind Meta AI — a strategic pivot away from Llama as Meta's frontier effort, though existing Llama models remain open and available for self-hosting via Ollama, Hugging Face, and every major inference framework.

Meta Superintelligence Labs Rising
Muse Spark 1.1 (current), Muse Spark 1.0
Core superpower
A multimodal reasoning model built for agentic tasks — planning, using tools, and orchestrating work across apps with a very large 1-million-token memory that it actively manages during long jobs.
Key trade-off
It's tuned for tool use and orchestration rather than raw coding accuracy, so it trails top rivals on the hardest coding and reasoning tasks — and all the standout benchmark numbers so far are Meta's own, not independent tests.
Speed
7.5/10
Reasoning
7.8/10
Cost
Low
Best non-technical use
Chatting free in the Meta AI app or at meta.ai, where "Thinking" mode shows the model's step-by-step reasoning before it answers — handy for multi-step questions like planning an event or working through a problem out loud.
Cost tier
Low

As of July 2026, the current model in this family is Muse Spark 1.1, which Meta released on July 9, 2026 as its second Meta Superintelligence Labs model and a significant upgrade over the original Muse Spark from April. It is a multimodal reasoning model built for agentic tasks, with a 1-million-token context window and gains in tool use, computer use, coding, and multimodal understanding, and now runs in "Thinking" mode inside the Meta AI app and at meta.ai. On the same day, Meta opened a public preview of the new Meta Model API — a self-serve, US-only developer preview priced at $1.25 per million input tokens and $4.25 per million output tokens, with $20 in free credits for new accounts.

DeepSeek Disruptor
DeepSeek V4 Pro · DeepSeek V4 Flash
Core superpower
Frontier-competitive reasoning at a small fraction of closed-model pricing — still the most significant price disruption in AI history
Key trade-off
Chinese-operated; some content restrictions; data privacy considerations for sensitive enterprise use
Speed
5.4/10
Reasoning
8.6/10
Cost
Ultra-low
Best non-technical use
High-volume automation where cost-per-prompt needs to be near zero; advanced coding tasks where quality needs to match closed frontier models at a fraction of the cost
Cost tier
Ultra-low

DeepSeek's January 2025 release of R1 shocked the AI industry by matching OpenAI o1's reasoning benchmark performance at a fraction of the training cost of Western frontier labs. DeepSeek V4 Pro is the current flagship, a Mixture of Experts model achieving frontier-competitive reasoning scores at dramatically lower inference cost than closed alternatives. DeepSeek V4 Flash trades some capability for speed and even lower cost. Both are available as open weights and via the DeepSeek API.

Alibaba Rising
Qwen3.8-Max-Preview, Qwen3.7-Max, Qwen3.7-Plus, Qwen3.6-27B (open weights)
Core superpower
A fast-shipping frontier lineup topped by a huge new preview model, backed by long-horizon agent skills and a genuinely open-weight budget tier you can download and run yourself.
Key trade-off
The lineup is split: the newest and strongest models (Qwen3.8-Max-Preview, Qwen3.7-Max, Qwen3.7-Plus) are cloud-only and proprietary, while the free-to-download open weights are a generation behind — so you can't have the top capability and self-hosting at the same time.
Speed
5/10
Reasoning
8.5/10
Cost
Mixed
Best non-technical use
Working through long, multi-step reasoning tasks — like analyzing a big pile of documents or research at once — thanks to a very large context window that fits huge amounts of text in a single request.
Cost tier
Low

Qwen is Alibaba's family of large language models. Its current flagship is Qwen3.8-Max, previewed on July 19, 2026 as a 2.4-trillion-parameter multimodal (text, image, video, document) sparse Mixture-of-Experts model , which Alibaba claims performs "second only to Fable 5" — though that claim shipped with no benchmark table, no model card, no license, and no active-parameter count, and nobody outside Alibaba can run it yet, so "open-weight" is a promise rather than a download. Below it sit the proprietary, API-only Qwen3.7-Max, formally announced at the Alibaba Cloud Summit in May 2026 and text-only , and Qwen 3.7 Plus, a low-cost multimodal sibling that adds vision and video and lists at roughly one-sixth the per-token price of Qwen 3.7 Max , both carrying a 1M-token context window. Note that Qwen 3.7-Max is not open source and not open-weight — you cannot download it or run it on your own GPUs ; if you specifically want to self-host, the genuinely open line is the earlier Qwen3.6 family, which remains a solid free option for cost-sensitive work.

Moonshot AI Rising
Kimi K3 (current flagship), K3 Swarm Max; prior: K2.7 Code, K2.6, K2.5, K2
Core superpower
A huge open-weight model that can hold an entire large project in view at once and reason across all of it without losing the thread.
Key trade-off
It always "thinks" before answering, and at launch you can't dial that down — so simple questions cost more time and money than they need to.
Speed
6/10
Reasoning
8.8/10
Cost
Mid
Best non-technical use
Working through a very long document set — say a stack of contracts or a full research folder — and getting answers that stay consistent from the first page to the last.
Cost tier
Low

Kimi K3, launched July 16, 2026, is Moonshot AI's current flagship and the successor to the entire Kimi K2 family (K2, K2.5, K2.6, and K2.7 Code). It is a roughly 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window, native vision, and always-on reasoning, positioned as the largest open-weight model released to date. Its standout use case is long-horizon work — coding across full codebases, analyzing large document sets, and running multi-step agent tasks — with two variants at launch, K3 Max for chat and agents and K3 Swarm Max for large-scale parallel processing.

Mistral AI Rising
Mistral Medium 3.5 · Mistral Large 3 · Mistral Small 4
Core superpower
Fast, efficient models with a genuinely strong price-to-performance ratio, and a unified Vibe agent now handling both research and coding tasks across web, IDE, and terminal
Key trade-off
The lineup has gotten genuinely complex — Medium 3.5 is the newest release and now the default in Mistral's own tools, but Large 3 remains the largest model for the heaviest workloads, so "which Mistral model" isn't a one-line answer anymore
Speed
9.5/10
Reasoning
7.1/10
Cost
Standard
Best non-technical use
Fast, cost-efficient production applications where European data sovereignty and open weights matter; real-time summarisation and classification at scale
Cost tier
Standard

Mistral AI, founded in France in 2023, has shipped a rapid succession of releases through 2026. Mistral Medium 3.5 is the newest, a 128-billion-parameter dense model released as open weights and now the default in Mistral's own Vibe and Le Chat tools. Mistral Large 3 remains the largest model in the lineup, a Mixture-of-Experts model released under Apache 2.0. Mistral Small 4 unifies the company's previously separate reasoning, multimodal, and coding models into a single configurable model. Mistral is a strong advocate for open-source AI and European digital sovereignty. Models are available via the Mistral API or as open weights for local deployment.