Skip to content

Context & Capacity Comparison

Compare context window sizes and output capacities across AI models. Larger context windows allow processing more text, code, or documents in a single request. Bar width uses a logarithmic scale.

40
Models Tracked
15
Providers
2M
Largest Context
Context Window
Max Output Tokens
Grok 4.20 Multi-Agent
xAIcoding
2M
Grok 4.20
xAIcoding
2M
Llama 4 Scout
Metacoding
1.3M(16K out)
GPT-5.6 Luna Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna
OpenAIcoding
1.1M(128K out)
GPT-5.6 Luna (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra
OpenAIcoding
1.1M(128K out)
GPT-5.6 Terra (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol Pro
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol
OpenAIcoding
1.1M(128K out)
GPT-5.6 Sol (batch)
OpenAIcoding
1.1M(128K out)
OpenAI GPT Latest
~openaicoding
1.1M(128K out)
GPT-5.5 Pro
OpenAIcoding
1.1M(128K out)
GPT-5.5 Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.5
OpenAIcoding
1.1M(128K out)
GPT-5.5 (batch)
OpenAIcoding
1.1M(128K out)
MiMo-V2.5-Pro
Xiaomicoding
1.1M(131K out)
MiMo-V2.5
Xiaomicoding
1.1M(131K out)
GPT-5.4 Pro
OpenAIcoding
1.1M(128K out)
GPT-5.4 Pro (batch)
OpenAIcoding
1.1M(128K out)
GPT-5.4
OpenAIcoding
1.1M(128K out)
GPT-5.4 (batch)
OpenAIcoding
1.1M(128K out)
LongCat 2.0
Meituancoding
1.0M(262K out)
Muse Spark 1.2
metacoding
1.0M
DeepSeek V4 Flash Latest
~deepseekcoding
1.0M(131K out)
DeepSeek V4 Flash 0731
DeepSeekcoding
1.0M(384K out)
Laguna S 2.1
poolsidecoding
1.0M(131K out)
Gemini 3.6 Flash
Googlecoding
1.0M(66K out)
Gemini 3.6 Flash (batch)
Googlecoding
1.0M(66K out)
Gemini 3.5 Flash Lite
Googlecoding
1.0M(66K out)
Gemini 3.5 Flash Lite (batch)
Googlecoding
1.0M(66K out)
Inkling
thinkingmachinescoding
1.0M(262K out)
Kimi K3
Moonshot AIcoding
1.0M
Muse Spark 1.1
metacoding
1.0M
GLM 5.2
Zhipu AIcoding
1.0M(128K out)
MiniMax M3
MiniMaxcoding
1.0M(512K out)

Context window sizes and output capacities are sourced from provider API data. The blue bar shows the full context window, while the darker bar shows the maximum output token limit. Scale is logarithmic to better visualize the range from thousands to millions of tokens.

Frequently Asked Questions

The context window is the total number of tokens a model can handle in one request, including both input and output. Max output tokens is the maximum length of the model's response alone. For example, a model with a 128K context window and 4K max output can accept up to 124K tokens of input but will only generate up to 4K tokens in its response.

Context window size depends on the model's architecture and training. Newer models like Gemini and Claude support over 1M tokens using advanced attention mechanisms. Older or smaller models may be limited to 4K-32K tokens. Larger contexts require more computational resources, which is why they often correlate with higher pricing.

Match the model to your task requirements. For document analysis or code review, prioritize large context windows. For content generation, prioritize high max output tokens. For chat applications, moderate values for both usually suffice. Also consider that not all models perform equally well at their maximum capacity, so test with your specific workload.

AI Model Context & Capacity Comparison - Size & Architecture | LM Market Cap