Skip to main content

LLM Statistics 2026: Key Numbers, Data & Facts

Current LLM statistics on models, providers, parameters, context windows, pricing, and benchmarks. Status September 2026.

FHFinn Hillebrandt
AI Technology
LLM Statistics 2026: Key Numbers, Data & Facts
Links marked with * are affiliate links. If a purchase is made through such links, we receive a commission.

Large language models are the heart of the AI revolution. But how many are there, really? Who builds them? What do they cost? And which model is actually the best?

The short answer:

It has gotten messy. In 2026, a new top-tier model shows up roughly every month, prices swing by a factor of 600, and the single most important metric of the past few years, the parameter count, is something the big labs no longer disclose at all.

In this article, I sort through the numbers. API models come from our centrally maintained LLM database, the same one that powers tools like the API cost calculator. Open weights live in the separately generated open source directory. Both sources reflect the state of September 2026.

TL;DRKey Takeaways
  • Our database tracks 288 LLMs from 35 providers, 126 of them proprietary and 162 openly available.
  • GPT-6 Astra is OpenAI's new flagship with a 1.05M-token context window, 128,000 output tokens, and an April 30, 2026 knowledge cutoff. It is available through the API and in ChatGPT Work and Codex for eligible plans.
  • Claude Opus 5.5 has been Anthropic's new Opus flagship since September 22, 2026: $4 / $20 per million tokens, 1M context, and a June 2026 knowledge cutoff. GPT-6 Sol and GPT-6 Luna round out OpenAI's lineup as cheaper successors to GPT-5.6 Sol and Luna, both with a 1.05M-token context window.
  • Qwen 3.8 Flash adds a multimodal API model with a native 1-million-token context. Qwen 3.8 Flash Next adds open weights with 125 billion main-model parameters, 51 billion N-gram embedding parameters, and 6 billion active parameters.
  • For coding, Claude Opus 5 still leads at 97.0% SWE-bench Verified per Vals AI, closely followed by the open DeepSeek V4 Pro at 96.4%. Vals AI archived the benchmark as saturated on September 1, 2026, so later models like Claude Opus 5.5 or GPT-6 Sol won't show up there.
  • Prices range from $0.05 (GPT-5 nano) to $30 (GPT-5.5 Pro) per 1M input tokens. Frontier labs no longer disclose parameter counts.

1. How Many Large Language Models Are There in 2026?

Our database currently tracks 288 large language models from 35 different providers, from GPT-2 in 2019 to the latest flagships in September 2026, including GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, Mythos 5.1, and Qwen 3.8 Flash. Qwen 3.8 Flash Next is listed separately in the open source directory so the open weights are not counted as a second API model. This is deliberately a curated selection of the most important models, not a claim to completeness.

For context:

According to the Stanford AI Index 2026, US labs alone shipped around 60 notable models in 2025, Chinese providers about 35. More than 90% of all significant frontier models now come from industry rather than academic research. The market has professionalized and concentrated.

2. The Biggest LLM Providers by Model Count

A simple indicator of how active a lab is: the number of models it maintains. The chart below shows how many of the models we track belong to each provider:

Source: gradually.ai LLM database
gradually.ai

Alibaba leads with 43 models, followed by OpenAI with 41 and Google with 29. That number only measures how deeply a lab maintains its lineup, though, not actual usage. Real market share looks different: in AI chatbot web traffic, ChatGPT dominates, while Gemini and Claude follow behind.

3. Parameters and Architecture: The End of Size Disclosures

For years, the parameter count was the most important metric for a model. GPT-3 had 175 billion, GPT-4 an estimated 1.76 trillion. Then the labs stopped reporting the number.

Today the rule is:

For every current frontier model from OpenAI, Anthropic, Google, and xAI, the parameter count is officially unknown. Model size has become a trade secret. Concrete, confirmed numbers almost only exist for open-weights models, and those are huge. The new leader is Moonshot AI's Kimi K3 at an official 2.8 trillion parameters (unveiled July 16, 2026, with weights released July 27 under Moonshot’s own “Kimi K3 License”):

Kimi K3 MoE, 896 experts (16 active)
2.8T
Qwen 3.8 Max 0902 MoE, 95B active
2.4T
DeepSeek-V4-Pro MoE, 49B active
1.6T
Kimi K2.6 MoE, 32B active
1T
GLM-5.2 MoE, 40B active
744B
DeepSeek V3.2 MoE, 37B active
685B
Mistral Large 3 MoE, 41B active
675B
DeepSeek-V4.1-Flash MoE, 8B active (input)
552B
Llama 4 Maverick MoE, 17B active
400B
Grok-1 MoE (2024)
314B
Source: gradually.ai LLM database
gradually.ai

The architecture is the striking part. Almost all large models today use a Mixture-of-Experts (MoE) design, where only a fraction of the parameters is active per request. DeepSeek-V4-Pro has 1.6 trillion parameters but activates only 49 billion per token, around 3%. Qwen 3.8 Flash Next combines 125 billion main-model parameters with 51 billion N-gram embedding parameters and activates 6 billion parameters per token. That makes giant models affordable to run. In total, 109 of the tracked models are built as MoE.

You can filter and search the full parameter database by provider, size, and type below. For most current frontier models, the parameter column deliberately reads "unknown":

Legend:

500B+
100-500B
20-100B
5-20B
Under 5B

Showing 288 models

Parameter sizes of popular Large Language Models (as of August 2026)
Model
Developer
Parameters
MiniMax M2.7
MoE
MiniMax
Unknown
MiniMax M2.5
MoE
MiniMax
Unknown
GLM-4.7
MoE
Z.ai
Unknown
GPT-6 Astra
OpenAI
Unknown
GPT-6 Sol
OpenAI
Unknown
GPT-6 Luna
OpenAI
Unknown
GPT-5.6 Sol
OpenAI
Unknown
GPT-5.6 Terra
OpenAI
Unknown
GPT-5.6 Luna
OpenAI
Unknown
GPT-5.5
OpenAI
Unknown
GPT-5.5 Pro
OpenAI
Unknown
GPT-5.5 Instant
OpenAI
Unknown
ChatGPT chat-latest
OpenAI
Unknown
GPT-5.4
OpenAI
Unknown
GPT-5.4 Pro
OpenAI
Unknown
GPT-5.4 mini
OpenAI
Unknown
GPT-5.4 nano
OpenAI
Unknown
GPT-5.3-Codex
OpenAI
Unknown
GPT-5.3 Instant
OpenAI
Unknown
GPT-5.2
OpenAI
Unknown
GPT-5.1 Instant
OpenAI
Unknown
GPT-5.1 Thinking
OpenAI
Unknown
GPT-5
OpenAI
Unknown
GPT-5 Pro
OpenAI
Unknown
GPT-5 mini
OpenAI
Unknown
GPT-5 nano
OpenAI
Unknown
GPT-4 Turbo
OpenAI
Unknown
GPT-4.1
OpenAI
Unknown
GPT-4.1 mini
OpenAI
Unknown
GPT-4.1 nano
OpenAI
Unknown
GPT-3.5 Turbo
OpenAI
Unknown
o3
OpenAI
Unknown
o3-pro
OpenAI
Unknown
o3-mini
OpenAI
Unknown
o4-mini
OpenAI
Unknown
o1
OpenAI
Unknown
o1-mini
OpenAI
Unknown
Claude Opus 5.5
Anthropic
Unknown
Claude Sonnet 5.5
Anthropic
Unknown
Claude Fable 5.1
Anthropic
Unknown
Claude Mythos 5.1
Anthropic
Unknown
Claude Fable 5
Anthropic
Unknown
Claude Mythos 5
Anthropic
Unknown
Claude Sonnet 5
Anthropic
Unknown
Claude Opus 5
Anthropic
Unknown
Claude Opus 4.8
Anthropic
Unknown
Claude Opus 4.7
Anthropic
Unknown
Claude Opus 4.6
Anthropic
Unknown
Claude Sonnet 4.6
Anthropic
Unknown
Claude Opus 4.5
Anthropic
Unknown
Claude Opus 4.1
Anthropic
Unknown
Claude Sonnet 4.5
Anthropic
Unknown
Claude Haiku 4.5
Anthropic
Unknown
Claude Sonnet 4
Anthropic
Unknown
Claude Opus 4
Anthropic
Unknown
Claude Sonnet 3.7
Anthropic
Unknown
Claude 3.5 Haiku
Anthropic
Unknown
Gemini 3.8 Flash
Google
Unknown
Gemini 3.8 Flash Cyber
Google
Unknown
Gemini 3.7 Flash
Google
Unknown
Gemini 3.6 Flash
Google
Unknown
Gemini 3.5 Flash
Google
Unknown
Gemini 3.1 Pro Preview
Google
Unknown
Gemini 3 Flash Preview
MoE
Google
Unknown
Gemini 3.5 Flash-Lite
Google
Unknown
Gemini 3.1 Flash-Lite
MoE
Google
Unknown
Gemini 2.5 Pro
MoE
Google
Unknown
Gemini 2.5 Flash
MoE
Google
Unknown
Gemini 2.5 Flash-Lite
MoE
Google
Unknown
Gemini 3 Pro
MoE
Google
Unknown
Gemini 2.0 Flash
MoE
Google
Unknown
Gemini 1.5 Pro
MoE
Google
Unknown
Grok 4.7
xAI
Unknown
Grok 4.6
xAI
Unknown
Grok 4.5
xAI
Unknown
Grok 4.3
xAI
Unknown
Grok 4.20 Reasoning
xAI
Unknown
Grok 4.20 Multi-Agent
xAI
Unknown
Grok Build 0.1
xAI
Unknown
Grok 4
xAI
Unknown
Grok 3
xAI
Unknown
Grok 2
xAI
Unknown
Qwen 3.7 Max
MoE
Alibaba
Unknown
Seed 1.8
ByteDance
Unknown
Seed 2.0 Pro
ByteDance
Unknown
Muse Spark 1.3
Meta
Unknown
Muse Spark 1.2
Meta
Unknown
KAT-Coder-Air V2.5
Kuaishou
Unknown
KAT-Coder-Pro V2.5
Kuaishou
Unknown
Seed 1.6 Flash
ByteDance
Unknown
Seed 2.0 Lite
ByteDance
Unknown
Seed 2.0 Mini
ByteDance
Unknown
Seed 2.0 Code
ByteDance
Unknown
Seed 2.1 Turbo
ByteDance
Unknown
Qwen 3.7 Flash
Alibaba
Unknown
Qwen 3.7 Plus
MoE
Alibaba
Unknown
Amazon Nova 2 Lite
Amazon
Unknown
Amazon Nova Premier
Amazon
Unknown
Amazon Nova Pro
Amazon
Unknown
Amazon Nova Lite
Amazon
Unknown
Amazon Nova Micro
Amazon
Unknown
Sonar
Perplexity
Unknown
Sonar Pro
Perplexity
Unknown
Sonar Reasoning Pro
Perplexity
Unknown
Sonar Deep Research
Perplexity
Unknown
Solar Pro 4
Upstage
Unknown
Kimi K3
MoE(104B active)
Moonshot AI
2.8T
Qwen 3.8 2.4T A95B
MoE(95B active)
Alibaba
2.4T
Qwen 3.8 Max 0902
MoE(95B active)
Alibaba
2.4T
Claude 3 Opus
Anthropic
2T*
Llama 4 Behemoth
MoE(288B active)
Meta
2T
GPT-4
MoE(220B active)
OpenAI
1.76T*
DeepSeek-V4-Pro
MoE(49B active)
DeepSeek
1.6T
LongCat 2.0
MoE(48B active)
Meituan
1.6T
Ring 2.6 1T
MoE(63B active)
inclusionAI
1T
Kimi K2.5
MoE(32B active)
Moonshot AI
1T
Kimi K2 Thinking
MoE(32B active)
Moonshot AI
1T
Kimi K2 0711
MoE(32B active)
Moonshot AI
1T
Kimi K2 0905
MoE(32B active)
Moonshot AI
1T
Kimi K2.6
MoE(32B active)
Moonshot AI
1T
Kimi K2.7 Code
MoE(32B active)
Moonshot AI
1T
Qwen 3.6 Max-Preview
MoE
Alibaba
1T*
Yi-Large
MoE
01.AI
1T
MiMo-V2.5-Pro
MoE(42B active)
Xiaomi
1T
MiMo-V2.5-Pro-UltraSpeed
MoE(42B active)
Xiaomi
1T
Inkling
MoE(41B active)
Thinking Machines Lab
975B
GLM-5.2
MoE(40B active)
Z.ai
744B
GLM-5.3
MoE(40B active)
Z.ai
744B
GLM-5.1
MoE(40B active)
Z.ai
744B
GLM-5
MoE(40B active)
Z.ai
744B
DeepSeek-V3 0324
MoE(37B active)
DeepSeek
685B
Mistral Large 3
MoE(41B active)
Mistral AI
675B
DeepSeek-V3.1 Terminus
MoE(37B active)
DeepSeek
671B
DeepSeek-V3.2 Exp
MoE(37B active)
DeepSeek
671B
DeepSeek-V3.1
MoE(37B active)
DeepSeek
671B
DeepSeek-R1 0528
MoE(37B active)
DeepSeek
671B
DeepSeek-V3.2
MoE(37B active)
DeepSeek
671B
DeepSeek-V3
MoE(37B active)
DeepSeek
671B
DeepSeek-R1
MoE(37B active)
DeepSeek
671B
DeepSeek-V4.1-Flash
MoE
DeepSeek
552B
Nemotron 3 Ultra
MoE(55B active)
NVIDIA
550B
PaLM
Google
540B
Megatron-Turing NLG
NVIDIA
530B
Qwen 3 Coder 480B A35B
MoE(35B active)
Alibaba
480B
MiniMax M1
MoE(45.9B active)
MiniMax
456B
MiniMax-01
MoE(45.9B active)
MiniMax
456B
MiniMax M3
MoE(23B active)
MiniMax
428B
ERNIE 4.5 VL 424B A47B
MoE(47B active)
Baidu
424B
Hermes 4 405B
Nous Research
405B
Llama 3.1 405B
Meta
405B
Llama 4 Maverick
MoE(17B active)
Meta
400B
Trinity Large Thinking
MoE(13B active)
Arcee AI
398B
Nex-N2-Pro
MoE(17B active)
Nex AGI
397B
Qwen 3.5 397B A17B
MoE(17B active)
Alibaba
397B
GLM-4.6
MoE(32B active)
Z.ai
355B
GLM-4.5
MoE(32B active)
Z.ai
355B
Nemotron-4 340B
NVIDIA
340B
PaLM 2
Google
340B*
GLM-5.3-Flash
MoE(18B active)
Z.ai
320B
Grok 1
MoE(86B active)
xAI
314B
MiMo-V2.5
MoE(15B active)
Xiaomi
310B
Hy3
MoE(21B active)
Tencent
295B
DeepSeek-V4-Flash
MoE(13B active)
DeepSeek
284B
Inkling Small
MoE(12B active)
Thinking Machines Lab
276B
DeepSeek-V2
MoE(21B active)
DeepSeek
236B
Qwen 3 235B A22B Instruct 2507
MoE(22B active)
Alibaba
235B
Qwen 3 235B A22B
MoE(22B active)
Alibaba
235B
Qwen 3 235B A22B Thinking 2507
MoE(22B active)
Alibaba
235B
Qwen 3 VL 235B A22B
MoE(22B active)
Alibaba
235B
Qwen 3 VL 235B A22B Thinking
MoE(22B active)
Alibaba
235B
MiniMax M2.1
MoE(10B active)
MiniMax
230B
MiniMax M2
MoE(10B active)
MiniMax
230B
Seed 1.6
MoE(23B active)
ByteDance
230B
GPT-4o
OpenAI
200B*
Step 3.7 Flash
MoE(11B active)
StepFun
198B
Step 3.5 Flash
MoE(11B active)
StepFun
196.8B
Falcon 180B
TII
180B
BLOOM
BigScience
176B
GPT-3
OpenAI
175B
Claude 3.5 Sonnet
Anthropic
175B*
OPT-175B
Meta
175B
Mixtral 8x22B
MoE(39B active)
Mistral AI
141B
LaMDA
Google
137B
DBRX
MoE(36B active)
Databricks
132B
Mistral Medium 3.5
Mistral AI
128B
Qwen 3.8 Flash
MoE(6B active)
Alibaba
125B
Ling 3.0 Flash
MoE(5.1B active)
inclusionAI
124B
Mistral Large 2
Mistral AI
123B
Qwen 3.5 122B A10B
MoE(10B active)
Alibaba
122B
Nemotron 3 Super
MoE(12B active)
NVIDIA
120B
Mistral Small 4
MoE(6.5B active)
Mistral AI
119B
Laguna S 2.1
MoE(8B active)
Poolside
118B
GPT OSS 120B
MoE(5.1B active)
OpenAI
117B
Command A
Cohere
111B
Llama 4 Scout
MoE(17B active)
Meta
109B
GLM-4.6V
MoE(12B active)
Z.ai
106B
GLM-4.5 Air
MoE(12B active)
Z.ai
106B
GLM-4.5V
MoE(12B active)
Z.ai
106B
Command R+
Cohere
104B
Solar Pro 3
MoE(12B active)
Upstage
102B
Hunyuan A13B Instruct
MoE(13B active)
Tencent
80B
Qwen 3 Coder Next
MoE(3B active)
Alibaba
80B
Qwen 3 Next 80B A3B Instruct
MoE(3B active)
Alibaba
80B
Qwen 3 Next 80B A3B Thinking
MoE(3B active)
Alibaba
80B
Qwen 2.5 72B
Alibaba
72.7B
Virtuoso Large
Arcee AI
72B
Qwen 2.5 VL 72B
Alibaba
72B
Hermes 4 70B
Nous Research
70B
DeepSeek-R1-Distill-Llama-70B
DeepSeek
70B
Claude 3 Sonnet
Anthropic
70B*
Llama 3.3 70B
Meta
70B
Llama 3.1 70B
Meta
70B
Llama 3 70B
Meta
70B
Llama 2 70B
Meta
70B
Mixtral 8x7B
MoE(12.9B active)
Mistral AI
46.7B
Falcon 40B
TII
40B
Nex-N2-mini
MoE(3B active)
Nex AGI
35B
Qwen 3.6 35B A3B
MoE(3B active)
Alibaba
35B
Qwen 3.5 35B A3B
MoE(3B active)
Alibaba
35B
Yi-34B
01.AI
34B
Laguna XS 2.1
MoE(3B active)
Poolside
33B
Qwen 3 32B
Alibaba
32.8B
Qwen 3 VL 32B
Alibaba
32.8B
Qwen 2.5 Coder 32B
Alibaba
32.5B
Qwen 2.5 32B
Alibaba
32B
Command R
Cohere
32B
Solar Pro 2
Upstage
31B
Gemma 4 31B
Google
30.7B
Qwen 3 30B A3B Instruct 2507
MoE(3.3B active)
Alibaba
30.5B
Qwen 3 Coder 30B A3B
MoE(3.3B active)
Alibaba
30.5B
Qwen 3 30B A3B
MoE(3.3B active)
Alibaba
30.5B
Qwen 3 30B A3B Thinking 2507
MoE(3.3B active)
Alibaba
30.5B
Qwen 3 VL 30B A3B
MoE(3.3B active)
Alibaba
30.5B
Qwen 3 VL 30B A3B Thinking
MoE(3.3B active)
Alibaba
30.5B
Nemotron 3 Nano
MoE(3.5B active)
NVIDIA
30B
Nemotron 3.5 Lightning
MoE(3B active)
NVIDIA
30B
GLM-4.7-Flash
MoE(3B active)
Z.ai
30B
Muse Glimmer 30B
Meta
29.6B
Gemma 3 27B
Google
27B
Qwen 3.8 27B
Alibaba
27B
Qwen 3.5 27B
Alibaba
27B
Qwen 3.6 27B
Alibaba
27B
Gemma 2 27B
Google
27B
Gemma 4 26B A4B
MoE(3.8B active)
Google
25.2B
Mistral Small 3.2
Mistral AI
24B
Mistral Small 3.1
Mistral AI
24B
Voxtral Small 24B
Mistral AI
24B
Mistral Small 3
Mistral AI
24B
GPT OSS 20B
MoE(3.6B active)
OpenAI
21B
Claude 3 Haiku
Anthropic
20B*
Qwen 3 14B
Alibaba
14.8B
Qwen 2.5 14B
Alibaba
14B
Phi-4
Microsoft
14B
Ministral 3 14B
Mistral AI
13.9B
Nemotron Nano 2 VL 12B
NVIDIA
12.6B
Gemma 3 12B
Google
12B
Mistral Nemo
Mistral AI
12B
Solar Mini
Upstage
10.7B
Qwen 3.5 9B
Alibaba
9B
Gemma 2 9B
Google
9B
Nemotron Nano 2 9B
NVIDIA
9B
Ministral 3 8B
Mistral AI
8.8B
Qwen 3 8B
Alibaba
8.2B
Qwen 3 VL 8B
Alibaba
8.2B
Qwen 3 VL 8B Thinking
Alibaba
8.2B
Granite 4.1 8B
IBM
8B
Gemma 3n E4B
Google
8B
GPT-4o mini
OpenAI
8B*
Llama 3.1 8B
Meta
8B
Llama 3 8B
Meta
8B
Ministral 8B
Mistral AI
8B
Qwen 2.5 7B
Alibaba
7.6B
Mistral 7B
Mistral AI
7B
Command R7B
Cohere
7B
Phi-4 Multimodal
Microsoft
5.6B
Gemma 3 4B
Google
4B
Ministral 3 3B
Mistral AI
3.8B
Phi-4 mini
Microsoft
3.8B
Phi-3 mini
Microsoft
3.8B
Gemini Nano 2
Google
3.3B
Llama 3.2 3B
Meta
3.2B
Granite 4.0 H Micro
IBM
3B
Ministral 3B
Mistral AI
3B
Gemma 2 2B
Google
2B
Gemini Nano 1
Google
1.8B
GPT-2
OpenAI
1.5B
Llama 3.2 1B
Meta
1.2B
Qwen 2.5 0.5B
Alibaba
0.5B

Parameter sizes of popular Large Language Models (as of August 2026)

4. Context Windows: From 200,000 to 10 Million Tokens

The context window determines how much text a model can process at once. Here the orders of magnitude have multiplied over the past two years. The overview below covers more than 320 current models, sortable and filterable by provider:

Legend:
1M+ Tokens
200K-1M Tokens
100K-200K Tokens
32K-100K Tokens
Under 32K Tokens
Showing 329 models
Context window sizes of current AI language models (as of August 2026)
Model
Developer
Context Window
Meta
10.5M
Alibaba
10M
Google
2M
Google
2M
xAI
2M
xAI
2M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
OpenAI
1.1M
Meta
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Google
1M
Moonshot AI
1M
Meta
1M
Meta
1M
Poolside
1M
OpenAI
1M
OpenAI
1M
OpenAI
1M
Google
1M
Google
1M
Google
1M
xAI
1M
xAI
1M
xAI
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
Anthropic
1M
DeepSeek
1M
DeepSeek
1M
DeepSeek
1M
MiniMax
1M
Alibaba
1M
Alibaba
1M
Alibaba
1M
Meituan
1M
Thinking Machines Lab
1M
Thinking Machines Lab
1M
Alibaba
1M
Alibaba
1M
Alibaba
1M
Alibaba
1M
Z.ai
1M
Z.ai
1M
Z.ai
1M
Xiaomi
1M
Xiaomi
1M
Xiaomi
1M
MiniMax
1M
NVIDIA
1M
NVIDIA
1M
NVIDIA
1M
NVIDIA
1M
Amazon
1M
Amazon
1M
Amazon
1M
MiniMax
1M
Upstage
512K
xAI
500K
xAI
500K
xAI
500K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
OpenAI
400K
Amazon
300K
Amazon
300K
Mistral AI
262.14K
Mistral AI
262.14K
Moonshot AI
262.14K
Moonshot AI
262.14K
Alibaba
262.14K
Alibaba
262.14K
Poolside
262.14K
Nex AGI
262.14K
Nex AGI
262.14K
Tencent
262.14K
Tencent
262.14K
inclusionAI
262.14K
inclusionAI
262.14K
Arcee AI
262.14K
Google
262.14K
Google
262.14K
Mistral AI
262.14K
Mistral AI
262.14K
Mistral AI
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Alibaba
262.14K
Moonshot AI
262.14K
Moonshot AI
262.14K
Moonshot AI
262.14K
xAI
256K
xAI
256K
xAI
256K
Mistral AI
256K
Mistral AI
256K
Alibaba
256K
ByteDance
256K
ByteDance
256K
Kuaishou
256K
Kuaishou
256K
ByteDance
256K
ByteDance
256K
ByteDance
256K
ByteDance
256K
ByteDance
256K
ByteDance
256K
StepFun
256K
StepFun
256K
Cohere
256K
Cohere
256K
AI21 Labs
256K
AI21 Labs
256K
AI21 Labs
256K
MiniMax
245.76K
MiniMax
204.8K
MiniMax
204.8K
MiniMax
204.8K
MiniMax
204.8K
Z.ai
204.8K
Z.ai
204.8K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
Anthropic
200K
OpenAI
200K
OpenAI
200K
OpenAI
200K
OpenAI
200K
OpenAI
200K
Z.ai
200K
Z.ai
200K
Perplexity
200K
Z.ai
200K
01.AI
200K
01.AI
200K
DeepSeek
163.84K
DeepSeek
163.84K
DeepSeek
163.84K
Meta
131.07K
Meta
131.07K
Meta
131.07K
Meta
131.07K
Meta
131.07K
xAI
131.07K
Mistral AI
131.07K
Alibaba
131.07K
Alibaba
131.07K
Alibaba
131.07K
Meta
131.07K
Upstage
131.07K
Baidu
131.07K
IBM
131.07K
IBM
131.07K
Arcee AI
131.07K
Nous Research
131.07K
Nous Research
131.07K
OpenAI
131.07K
OpenAI
131.07K
Mistral AI
131.07K
Mistral AI
131.07K
Alibaba
131.07K
Alibaba
131.07K
Alibaba
131.07K
Alibaba
131.07K
Alibaba
131.07K
Moonshot AI
131.07K
Z.ai
131.07K
Z.ai
131.07K
Z.ai
131.07K
Meta
128K
Meta
128K
Meta
128K
Google
128K
Google
128K
Google
128K
xAI
128K
OpenAI
128K
OpenAI
128K
OpenAI
128K
OpenAI
128K
OpenAI
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
DeepSeek
128K
Mistral AI
128K
Mistral AI
128K
Mistral AI
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
Alibaba
128K
NVIDIA
128K
NVIDIA
128K
Cohere
128K
Perplexity
128K
Perplexity
128K
Perplexity
128K
DeepSeek
128K
DeepSeek
128K
Alibaba
128K
Cohere
128K
Cohere
128K
Amazon
128K
Microsoft
128K
Microsoft
128K
Microsoft
128K
Microsoft
128K
Microsoft
128K
Microsoft
128K
01.AI
128K
01.AI
128K
Nvidia
128K
Nvidia
128K
Nvidia
128K
Reka
128K
Reka
128K
Reka
128K
Zhipu AI
128K
Zhipu AI
128K
Baidu
128K
Mistral AI
65.54K
Upstage
65.54K
Z.ai
65.54K
Microsoft
64K
Mistral AI
32.77K
Mistral AI
32.77K
Mistral AI
32.77K
Alibaba
32.77K
Alibaba
32.77K
Alibaba
32.77K
Upstage
32.77K
Google
32.77K
Microsoft
32.77K
Databricks
32.77K
Google
32K
Mistral AI
32K
01.AI
32K
Microsoft
16.38K
01.AI
16K
Google
8.19K
Google
8.19K
OpenAI
8.19K
AI21 Labs
8.19K
Zhipu AI
8.19K
Baidu
8K
Cohere
4.1K
Nvidia
4.1K
Stability AI
4.1K
Stability AI
4.1K

Context window sizes of current AI language models (as of August 2026)

At the top are Llama 4 Scout and Qwen-Long with 10 million tokens each. That's roughly 30 Harry Potter books in a single prompt. Current all-rounders mostly sit around 1 million tokens, including GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, Claude Opus 5.5, Fable 5.1 and Mythos 5.1, Claude Opus 5, GPT-5.5, Gemini 3.8 Flash, Muse Spark 1.3, Qwen 3.8 Flash, and Qwen 3.8 Max 0902. Qwen 3.8 Flash Next offers a native 262,144-token window and can be extended to 1 million tokens with YaRN. For more on the individual model families, see our overviews of the Claude models and Gemini models.

5. What Does an LLM Cost? Prices per 1 Million Tokens

API prices span worlds. The cheapest model with API access is GPT-5 nano at $0.05 per 1M input tokens. The most expensive is GPT-5.5 Pro at $30, a 600x difference.

More interesting than the raw price is the ratio of price to performance. The chart below plots input price against coding performance (SWE-bench Verified). Models toward the bottom right are ideal: strong and cheap.

Price-performance: SWE-bench vs. input price
Anthropic
OpenAI
Google
DeepSeek
Moonshot AI
Efficiency frontier (best price-performance)
Source: gradually.ai LLM database
gradually.ai

The quiet star of this chart is DeepSeek-V4-Pro. Even the April preview build scores 77.4% on SWE-bench at Vals AI and, at just $0.66 input price, sits right on the efficiency frontier, no other model is both stronger and cheaper. The generally available build from August 13 reaches 96.4%. So if you don't strictly need the last few percentage points of coding performance, the open models offer an extremely good price-performance ratio. Qwen 3.8 Flash is listed at $0.113 input and $0.382 output per 1 million tokens in the official price list. QwenCloud uses the same API name, qwen3.8-flash, for Qwen 3.8 Flash Next, so the price is not counted as a second row. For a detailed cost estimate of your specific usage, see the API cost calculator.

6. LLM Performance Head to Head

To make the strengths and weaknesses of the top models visible at a glance, the radar below compares five representative frontier models across four dimensions: reasoning, coding, context window, and price efficiency. Each axis is scaled relative to the five models so even small leads become visible. The real values appear in the tooltip.

Claude Opus 5
Gemini 3.1 Pro
Gemini 3.5 Flash
Claude Sonnet 4.6
GPT-5.5
Sources: Artificial Analysis, Vals AI
gradually.ai

The pattern is clear. Claude Opus 5 and GPT-5.5 dominate on raw coding performance but are expensive. Gemini 3.5 Flash flips that, nearly on par on reasoning and only trailing on coding, yet with the best price efficiency in the field. Every AI project comes down to this one trade-off in the end, maximum quality versus maximum economy.

7. Open Source vs. Proprietary

One of the most important developments of 2026 is the catch-up of open models. Of the 288 tracked models, 126 are proprietary and 162 are openly available, 157 of them open-weights and 5 fully open-source.

But at the very top:

According to the Stanford AI Index 2026, the best closed model led the best open-weights model by 3.3 percentage points in early 2026. In August 2024, the gap had been only 0.5 percentage points. So at the top it has not been shrinking but widening again, with six of the top-ten models in the Chatbot Arena now closed once more. Coding looks different. In the Vals AI harness, the generally available DeepSeek-V4-Pro build from August 13, 2026 scores 96.4% on SWE-bench Verified, just 0.6 percentage points behind Claude Opus 5 at 97.0%. Vals AI archived the benchmark as saturated on September 1, 2026, so the top scores barely separate the models anymore. For an overview of the best free models, see our article on open-source LLMs.

How the license mix breaks down by provider is shown below: column width represents the number of tracked models per provider, and the colors mark the license type.

14%86%Alibaba4393%OpenAI4169%31%Google29100%Anthropic2488%Mistral AI1713%81%Meta16
Proprietary
Open-weights
Open-source
Source: gradually.ai LLM database
gradually.ai

8. Knowledge Cutoff: How Current Are the Models?

Every model has a knowledge cutoff, after which it has learned nothing more about the world. Claude Opus 5.5 currently has the freshest cutoff in our database, June 2026, closely followed by GPT-6 Luna at May 18, 2026:

Claude Opus 5.5
June 2026
Claude Fable 5.1
June 2026
Claude Mythos 5.1
June 2026
GPT-6 Luna
May 2026
GPT-6 Astra
Apr. 2026
GPT-6 Sol
Apr. 2026
Claude Fable 5
Jan. 2026
Claude Opus 4.8
Jan. 2026
GPT-5.5
Dec. 2025
GPT-5.5 Instant
Dec. 2025
GPT-5.3 Codex
Aug. 2025
GPT-5.2
Aug. 2025
Claude Opus 4.6
May 2025
Claude Sonnet 4.6
May 2025
Claude Opus 4.5
Mar. 2025
Gemini 3.1 Pro
Jan. 2025
Gemini 3 Flash
Jan. 2025
Gemini 2.5 Pro
Jan. 2025
DeepSeek R1
Jan. 2025
DeepSeek V3.1
Dec. 2024
Grok 4.1
Nov. 2024
Qwen3-Max
Nov. 2024
GPT-5
Oct. 2024
Llama 4 Scout
Aug. 2024
Gemini 2.0 Flash
Aug. 2024
Amazon Nova Pro
Aug. 2024
GPT-4.1
June 2024
GPT-5 mini
May 2024
Source: gradually.ai LLM database
gradually.ai

Between the knowledge cutoff and the release date there are usually six to eight months in which the model is trained and tested. For current events, the models therefore almost always need a web search. Raw model knowledge is always a few months old.

9. Release Pace: The Cadence of the Labs

How fast the market moves shows in the release timeline. What happened quarterly in 2024 comes almost monthly in 2026:

May 2024
GPT-4o
OpenAI makes real-time multimodal models the default.
Jan. 2025
DeepSeek-R1
First open reasoning model at frontier level, kicking off the open-weights wave.
June 2025
GPT-5
OpenAI merges reasoning and standard mode into one model family.
Dec. 2025
Gemini 3 Pro
Google opens the third Gemini generation with its first model.
Dec. 2025
GPT-5.2
OpenAI follows up with an improved reasoning update.
Dec. 2025
Mistral Large 3
Mistral counters with an open MoE model from Europe.
Feb. 2026
Claude Opus 4.6
Anthropic raises the reasoning bar with the new Opus.
Feb. 2026
Gemini 3.1 Pro
Google takes the GPQA Diamond lead at 94.3%.
April 2026
GPT-5.5
Scores 82.6% SWE-bench Verified in the Vals AI harness.
April 2026
Claude Opus 4.7
Anthropic reaches 82.0% on coding, just behind GPT-5.5.
April 2026
DeepSeek-V4-Pro
Open model hits 80.6% SWE-bench at a fraction of the price.
May 2026
Claude Opus 4.8
Hits 88.6% SWE-bench, the active coding benchmark at the time.
May 2026
Gemini 3.5 Flash
Google ships a fast, price-efficient Flash model.
June 2026
Claude Fable 5
Anthropic expands the lineup with a specialized variant. Back online since July 1 after a June 12-30 export-control pause, now the new coding benchmark at 95.0% SWE-bench Verified.
June 2026
Claude Mythos 5
A second specialized model, available through the API at first.
July 2026
GPT-5.6 Sol/Terra/Luna GA
OpenAI makes the GPT-5.6 family generally available on July 9, according to Axios/Bloomberg one day after government restrictions were lifted. Sol default in Codex (Ultra mode there from Plus), Terra free for Free/Go only inside Codex so far, while GPT-5.5 Instant remains the regular chat default. Sol at 88.8% on Terminal-Bench 2.1 (Sol Ultra 91.9%), 64.6% on SWE-Bench Pro, 80 on the Coding Agent Index.
July 2026
Kimi K3
Moonshot AI unveils its 2.8-trillion-parameter MoE on July 16: 1M context, native vision, 93.5% on GPQA Diamond, API pricing $3/$15. Weights have been available since July 27 under Moonshot’s own “Kimi K3 License”, making it the largest open-weight model ever released.
July 2026
Gemini 3.6 Flash
Google ships the new, more efficient Flash generation on July 21: 17% fewer output tokens than the still-active 3.5 Flash, $7.50 instead of $9 per 1M output tokens, March 2026 knowledge cutoff. Gemini 3.5 Flash-Lite also launches for high-throughput tasks.
July 2026
Claude Opus 5
Anthropic releases Opus 5 on July 24, close to Fable 5 in performance but at half the price ($5/$25) and with a May 2026 knowledge cutoff. Opus 4.8 is now considered superseded.
August 2026
Grok 4.6
xAI releases its new flagship for coding and long-running agents on August 12: 500K context, February 2026 knowledge cutoff, and pricing from $2 / $6 per 1M tokens.
August 2026
Gemini 3.7 Flash
Google releases a stable multimodal Flash model on August 13 with a 1M context window, 64K output, and introductory pricing of $0.75 / $3.75 through the end of 2026.
September 2026
Claude Fable 5.1 and Mythos 5.1
Anthropic releases its new frontier tier with a 1M context window, 128K output, a June 2026 knowledge cutoff, and cache reads at one quarter of the previous price.
September 2026
GPT-6 Astra
OpenAI begins a phased rollout of its new flagship on September 3: a 1.05M token context window, 128,000 output tokens, an April 30, 2026 knowledge cutoff, and it appears as GPT-6 Pro in ChatGPT.
September 2026
Gemini 3.8 Flash
Google releases a stable multimodal Flash model with a 1M context window, 64K output, and stronger coding and agentic performance.
September 2026
Muse Spark 1.3
Meta improves long-running agentic and coding tasks while reporting about 20% fewer tool calls and about 25% fewer tokens than version 1.2.
September 2026
Qwen 3.8 Max 0902
Alibaba updates its 2.4-trillion-parameter flagship for coding, long-horizon autonomous tasks, agent collaboration, and visual analysis.
September 2026
DeepSeek-V4.1-Flash
DeepSeek releases the MoE model V4.1-Flash on September 10: 552 billion parameters, with 8 billion active on input and 16 billion on output, plus native vision. The previous V4-Flash is retired.
September 2026
Grok 4.7
xAI releases its most capable model yet for coding and knowledge work on September 21, at the same $2 / $6 per million token pricing as Grok 4.6.
September 2026
Claude Opus 5.5
Anthropic releases its new Opus flagship on September 22: $4 / $20 per million tokens, 1M context, a June 2026 knowledge cutoff, and per Anthropic, Fable 5.1-level performance at 40% less cost than Opus 5.
September 2026
GPT-6 Sol and GPT-6 Luna
OpenAI adds two cheaper GPT-6 variants on September 22: Sol at $2 / $10, succeeding GPT-5.6 Sol, and Luna at $0.10 / $0.50, succeeding GPT-5.6 Luna, both with a 1.05M token context window.

Plotting every tracked model onto its release month makes the clustering visible: the darker a cell, the more models shipped that month.

JanFebMarAprMayJunJulAugSepOctNovDec2019120201202120222111120231212332024115534718319202586614311131111841420263181115816151213
Releases: lowhigh(max 18)
Source: gradually.ai LLM database
gradually.ai

December 2025 was especially dense, when Google, OpenAI, and Mistral all shipped new flagships in the same month. So was April 2026, which brought GPT-5.5, Claude Opus 4.7, DeepSeek-V4-Pro, Kimi K2.6, and Qwen 3.6 Max, five top models at once. Qwen 3.8 Flash and Qwen 3.8 Flash Next followed at the end of August 2026, and September 2026 piled on again with GPT-6 Astra, Claude Fable 5.1, Mythos 5.1, and, on September 22, three more models at once (Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna). If you want to keep up here, don't cling too tightly to individual version numbers.

10. Model Status: Active, Deprecated, Legacy

Not every model ever released is still usable. Across the three big providers Anthropic, Google, and OpenAI, we track the lifecycle of 97 models. Here is how they split across the individual statuses:

97models
Active3940.2%
Deprecated3839.2%
Legacy77.2%
Limited access44.1%
Pro-exclusive22.1%
API only22.1%
Preview33.1%
Open source22.1%
Source: gradually.ai LLM database
gradually.ai

Just over half of the models are still active, and nearly a third are already deprecated. And lifecycles are getting shorter. A good example is Gemini 3 Pro, deprecated only about three months after its release because Gemini 3.1 Pro was already standing by as a successor. Anyone building production systems on a model has to keep an active eye on these deprecations.

11. Market Position and Conclusion

The LLM market of 2026 has grown up. Instead of one dominant model, there's a tight leading pack of OpenAI, Anthropic, and Google, closely chased by open models from China, led by DeepSeek and Moonshot.

Bottom line:

Performance at the top is remarkably close together, and the competition is shifting to price, context length, and specialization. For most applications in 2026, it matters less which model is the absolute best and more which one is right for the specific purpose and budget. If you want to dig deeper into individual providers, you'll find the details in our statistics on OpenAI, Anthropic, Google Gemini, Grok, and DeepSeek.

Frequently Asked Questions

FH

Finn Hillebrandt

AI Expert & Blogger

Finn Hillebrandt is the founder of Gradually AI, an SEO and AI expert. He helps online entrepreneurs simplify and automate their processes and marketing with AI. Finn shares his knowledge here on the blog in 50+ articles as well as through the AI Business Club.

Learn more about Finn and the team, follow Finn on LinkedIn, join his Facebook group for ChatGPT, OpenAI & AI Tools or do like 17,500+ others and subscribe to his AI Newsletter with tips, news and offers about AI tools and online business. Also visit his other blog, Blogmojo, which is about WordPress, blogging and SEO.