Skip to content

Context Window Efficiency Explorer

Analyzes score-per-context-token ratio across 300 AI models to find those that make the best use of their context window, output capacity, and cost.

Context Window vs Score

LMMarketCap.com

Efficiency Overview

Key efficiency metrics across all analyzed models.

Most Efficient (128K+)

GPT-5.2 Chat

707.0 score/MToken

Best Output Efficiency

Gemma 2 27B

37.8 score/1K output

Best Cost Efficiency

Ling 3.0 Flash

952.4 score/$

Avg Overall Efficiency

6.3%

normalized across all models

Efficiency Rankings

Top 50 models ranked by score per million context tokens.

#ModelScoreContextScore/MToken
1Gemma 2 27BGoogle778K9448.2
2GPT-4OpenAI658K7923.3
3R1 Distill Llama 70BDeepSeek418K4968.3
4Hy-MT2-1.8BTencent408K4882.8
5Hy-MT2-30B-A3BTencent408K4882.8
6Hy-MT2-7BTencent408K4882.8
7Phi 4Microsoft6016K3674.3
8GLM 5.2 (free)Zhipu AI7633K2310.2
9R1DeepSeek7464K1153.1
10MiniMax M2-herMiniMax6866K1039.1
11Mixtral 8x22B InstructMistral AI6366K967.4
12GLM 4.5VZhipu AI6266K950.6
13Qwen3 30B A3B Thinking 2507Alibaba6482K782.5
14GPT-5.2 ChatOpenAI91128K707.0
15LFM2.5-2.6B (free)Liquid AI4066K610.4
16GLM 4.5Zhipu AI75131K573.0
17GPT-4o (batch)OpenAI72128K562.5
18GPT-4o (2024-11-20)OpenAI71128K556.3
19GPT-4o (2024-08-06)OpenAI71128K556.3
20GPT-4oOpenAI71128K556.3
21GPT-4o (2024-05-13)OpenAI71128K556.3
22GPT-4o-miniOpenAI69128K541.4
23GLM 4.5 AirZhipu AI71131K539.4
24Qwen3 VL 235B A22B ThinkingAlibaba69131K528.7
25GPT-4 TurboOpenAI67128K521.1
26Mistral LargeMistral AI66128K514.8
27Llama 3.3 70B InstructMeta67131K509.6
28DeepSeek V3.2DeepSeek83164K509.0
29Mistral Large 2407Mistral AI66131K502.8
30GLM 4.6VZhipu AI66131K500.5
31Qwen3 235B A22B Thinking 2507Alibaba65131K498.2
32Llama 3.1 70B InstructMeta65131K498.2
33Qwen3 30B A3BAlibaba64131K489.0
34GPT-4o-mini (batch)OpenAI63128K488.3
35R1 0528DeepSeek79164K484.6
36GPT-4 Turbo (batch)OpenAI62128K480.5
37Mercury 2Inception61128K475.8
38gpt-oss-120bOpenAI62131K470.0
39Qwen3 8BAlibaba61131K465.4
40GPT-4o-mini (2024-07-18)OpenAI57128K441.4
41DeepSeek V3.2 ExpDeepSeek72164K438.2
42DeepSeek V3.1DeepSeek72164K438.2
43DeepSeek V3 0324DeepSeek72164K438.2
44gpt-oss-20bOpenAI57131K437.9
45o3 ProOpenAI87200K433.5
46o3OpenAI87200K433.5
47o3 (batch)OpenAI87200K433.5
48Claude Opus 4.5Anthropic85200K425.5
49Claude Opus 4.5 (batch)Anthropic85200K425.5
50DeepSeek V3DeepSeek70164K424.2

Tier Analysis

Efficiency breakdown across context window tiers.

Small7 models
Avg Score52
Score/MToken5808.9

Worst

Phi 4

Medium7 models
Avg Score64
Score/MToken1116.2
Large148 models
Avg Score67
Score/MToken320.7
Mega138 models
Avg Score70
Score/MToken67.0

Diminishing Returns Analysis

Are bigger context windows correlated with higher scores?

TierAvg ContextAvg ScoreAvg Efficiency
Small9K525808.9
Medium63K641116.2
Large239K67320.7
Mega1.1M7067.0

Output Token Efficiency

Top 20 models by output efficiency (score per 1K output tokens). Models with 16K+ output tokens are highlighted.

ModelScoreMax OutputOutput Eff.
Gemma 2 27BGoogle772K37.8
MiniMax M2-herMiniMax682K33.3
GPT-4o (2024-05-13)OpenAI714K17.4
GPT-4 TurboOpenAI674K16.3
GPT-4OpenAI654K15.8
GPT-4 Turbo (batch)OpenAI624K15.0
Claude 3 HaikuAnthropic514K12.5
Command R (08-2024)Cohere494K12.2
Command R+ (08-2024)Cohere494K12.2
Schematron V2 Smallinference-net404K9.8
Hy-MT2-1.8BTencent404K9.8
Hy-MT2-30B-A3BTencent404K9.8
Hy-MT2-7BTencent404K9.8
Qwen3 8BAlibaba618K7.4
Qwen3 235B A22BAlibaba548K6.6
Command ACohere518K6.2
R1 Distill Llama 70BDeepSeek417K5.5
Gemma 4 31BGoogle16K+8116K4.9
Schematron V2 Turboinference-net408K4.9
LFM2.5-2.6B (free)Liquid AI408K4.9

Key Insights

Auto-generated observations from the efficiency data.

Context Sweet Spot

Small models have the highest average efficiency at 5808.9 score/MToken across 7 models.

Output Matters

Models with 16K+ output tokens score 30% higher on average than models with smaller output limits.

Compact High Performers

0 models achieve top-20 scores with under 128K context.

Explore More

Dive deeper into context windows, compare models, or explore other dimensions.

Frequently Asked Questions

Efficiency is measured as the score-per-context-token ratio - how much ranking score a model achieves relative to its context window size. Models that score highly with smaller context windows are considered more efficient than those requiring massive context to achieve similar results.

Cost efficiency combines quality (composite score) with pricing. The most cost-efficient models achieve high benchmark scores while maintaining low per-token API costs. Free and budget-tier models that perform well are the most cost-efficient options.

Not necessarily. Our efficiency analysis shows diminishing returns beyond certain context sizes. Models with 128K tokens often score similarly to those with 1M+ tokens, meaning the extra context capacity adds cost without proportional quality gains for most use cases.

AI Model Efficiency Explorer - Score Per Context Analysis | LM Market Cap