Skip to content

Context & Capacity Comparison

Compare context window sizes and output capacities across AI models. Larger context windows allow processing more text, code, or documents in a single request. Bar width uses a logarithmic scale.

40
Models Tracked
8
Providers
2M
Largest Context
Context Window
Max Output Tokens
Grok 4.20 Multi-Agent
xAIcoding
2M(1.8M out)
Grok 4.20
xAIcoding
2M(1.8M out)
GLM Flash Latest
~z-aicoding
1.3M(944K out)
GLM 5.3 Flash
Zhipu AIcoding
1.3M(944K out)
GLM Latest
~z-aicoding
1.3M(131K out)
GLM 5.3
Zhipu AIcoding
1.3M(131K out)
DeepSeek V4 Flash Latest
~deepseekcoding
1.3M(944K out)
DeepSeek V4 Flash 0731
DeepSeekcoding
1.3M(944K out)
Llama 4 Scout
Metacoding
1.3M(16K out)
GPT-6 Luna Pro
OpenAIcoding
1.1M(128K out)
GPT-6 Luna Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-6 Luna
OpenAIcoding
1.1M(128K out)
GPT-6 Luna (batch)
OpenAIcoding
1.1M(128K out)
GPT-6 Sol Pro
OpenAIcoding
1.1M(128K out)
GPT-6 Sol Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-6 Sol
OpenAIcoding
1.1M(128K out)
GPT-6 Sol (batch)
OpenAIcoding
1.1M(128K out)
GPT Astra Latest
~openaicoding
1.1M(128K out)
GPT Sol Latest
~openaicoding
1.1M(128K out)
GPT Terra Latest
~openaicoding
1.1M(128K out)
GPT Luna Latest
~openaicoding
1.1M(128K out)
GPT-6 Astra
OpenAIcoding
1.1M(128K out)
GPT-6 Astra (batch)
OpenAIcoding
1.1M(128K out)
GPT-6 Astra Pro
OpenAIcoding
1.1M(128K out)
GPT-6 Astra Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.5 Pro
OpenAIcoding
1.1M(128K out)
GPT-5.5 Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.5
OpenAIcoding
1.1M(128K out)

Context window sizes and output capacities are sourced from provider API data. The blue bar shows the full context window, while the darker bar shows the maximum output token limit. Scale is logarithmic to better visualize the range from thousands to millions of tokens.

Frequently Asked Questions

The context window is the total number of tokens a model can handle in one request, including both input and output. Max output tokens is the maximum length of the model's response alone. For example, a model with a 128K context window and 4K max output can accept up to 124K tokens of input but will only generate up to 4K tokens in its response.

Context window size depends on the model's architecture and training. Newer models like Gemini and Claude support over 1M tokens using advanced attention mechanisms. Older or smaller models may be limited to 4K-32K tokens. Larger contexts require more computational resources, which is why they often correlate with higher pricing.

Match the model to your task requirements. For document analysis or code review, prioritize large context windows. For content generation, prioritize high max output tokens. For chat applications, moderate values for both usually suffice. Also consider that not all models perform equally well at their maximum capacity, so test with your specific workload.

AI Model Context & Capacity Comparison - Size & Architecture | LM Market Cap