Skip to content

上下文窗口效率探索器

分析300个AI模型的评分/上下文令牌比率,找出最充分利用上下文窗口、输出容量和成本的模型。

上下文窗口 vs 评分

LMMarketCap.com

效率概览

所有分析模型的关键效率指标。

最高效 (128K+)

GPT-5.3 Chat

707.0 score/MToken

最佳输出效率

Gemma 2 27B

37.8 score/1K output

最佳成本效率

Ling-2.6-flash

2000.0 score/$

平均总体效率

4.4%

所有模型标准化

效率排名

按每百万上下文令牌评分排名的前50个模型。

#模型评分上下文评分/百万令牌
1Gemma 2 27BGoogle778K9448.2
2GPT-4OpenAI658K7911.1
3R1 Distill Llama 70BDeepSeek418K4968.3
4Phi 4Microsoft6016K3674.3
5Reka Edgerekaai4016K2441.4
6Perceptron Mk1perceptron4033K1220.7
7MiniMax M2-herMiniMax6966K1054.4
8Mixtral 8x22B InstructMistral AI6366K967.4
9GLM 4.5VZhipu AI6366K953.7
10Olmo 3 32B ThinkAllen AI5566K837.7
11Qwen3 30B A3B Thinking 2507Alibaba6482K782.5
12GPT-5.3 ChatOpenAI91128K707.0
13GPT-5.2 ChatOpenAI91128K707.0
14GLM 4.5Zhipu AI75131K573.0
15GPT-4o (batch)OpenAI72128K562.5
16GPT-4o (2024-11-20)OpenAI71128K556.3
17GPT-4o (2024-08-06)OpenAI71128K556.3
18GPT-4oOpenAI71128K556.3
19GPT-4o (2024-05-13)OpenAI71128K556.3
20GPT-4o-miniOpenAI69128K541.4
21GLM 4.5 AirZhipu AI71131K539.4
22Qwen3 VL 235B A22B ThinkingAlibaba69131K529.5
23GPT-4 TurboOpenAI67128K521.1
24Mistral LargeMistral AI66128K514.8
25Llama 3.3 70B InstructMeta67131K509.6
26GPT-4 Turbo PreviewOpenAI65128K506.2
27Mistral Large 2407Mistral AI66131K502.8
28GLM 4.6VZhipu AI66131K499.7
29Llama 3.1 70B InstructMeta65131K498.2
30DeepSeek V3.2DeepSeek81164K496.2
31Qwen3 30B A3BAlibaba64131K489.0
32GPT-4o-mini (batch)OpenAI63128K488.3
33R1 0528DeepSeek79164K484.6
34GPT-4 Turbo (batch)OpenAI62128K480.5
35Mercury 2Inception61128K476.6
36Qwen3 8BAlibaba61131K465.4
37R1DeepSeek74164K452.9
38GPT-4o-mini (2024-07-18)OpenAI57128K441.4
39DeepSeek V3.2 ExpDeepSeek72164K438.2
40DeepSeek V3.1DeepSeek72164K438.2
41DeepSeek V3 0324DeepSeek72164K438.2
42gpt-oss-20bOpenAI57131K437.9
43gpt-oss-20b (free)OpenAI57131K437.9
44o3 ProOpenAI87200K433.5
45o3 Pro (batch)OpenAI87200K433.5
46o3OpenAI87200K433.5
47o3 (batch)OpenAI87200K433.5
48Claude Opus 4.5Anthropic85200K425.5
49Claude Opus 4.5 (batch)Anthropic85200K425.5
50DeepSeek V3DeepSeek70164K424.2

层级分析

不同上下文窗口层级的效率分析。

Small5 models
平均评分57
评分/百万令牌5688.7

最差

Reka Edge

Medium6 models
平均评分59
评分/百万令牌969.4
Large174 models
平均评分65
评分/百万令牌311.8
Mega115 models
平均评分73
评分/百万令牌73.1

边际递减分析

更大的上下文窗口是否与更高的评分相关?

层级平均上下文平均评分平均效率
Small11K575688.7
Medium63K59969.4
Large236K65311.8
Mega1.0M7373.1

输出令牌效率

按输出效率(每1K输出令牌评分)排名的前20个模型。16K+输出令牌的模型已高亮显示。

模型评分最大输出输出效率
Gemma 2 27BGoogle772K37.8
MiniMax M2-herMiniMax692K33.7
GPT-4o (2024-05-13)OpenAI714K17.4
GPT-4 TurboOpenAI674K16.3
GPT-4 Turbo PreviewOpenAI654K15.8
GPT-4OpenAI654K15.8
GPT-4 Turbo (batch)OpenAI624K15.0
Claude 3 HaikuAnthropic514K12.5
Command R (08-2024)Cohere494K12.2
Command R+ (08-2024)Cohere494K12.2
Qwen3 8BAlibaba618K7.4
Qwen3 235B A22BAlibaba548K6.6
Command ACohere518K6.2
GPT-5.3 ChatOpenAI16K+9116K5.5
GPT-5.2 ChatOpenAI16K+9116K5.5
R1 Distill Llama 70BDeepSeek418K5.0
Nemotron 3.5 Content Safety (free)NVIDIA408K4.9
Perceptron Mk1perceptron408K4.9
Palmyra X5Writer408K4.9
Falcon-H1-Arabic 34B InstructTII408K4.9

关键洞察

从效率数据中自动生成的观察结果。

上下文最优点

Small models have the highest average efficiency at 5688.7 score/MToken across 5 models.

输出很重要

Models with 16K+ output tokens score 29% higher on average than models with smaller output limits.

紧凑型高性能模型

0 models achieve top-20 scores with under 128K context.

探索更多

深入了解上下文窗口、对比模型或探索其他维度。

Frequently Asked Questions

Efficiency is measured as the score-per-context-token ratio - how much ranking score a model achieves relative to its context window size. Models that score highly with smaller context windows are considered more efficient than those requiring massive context to achieve similar results.

Cost efficiency combines quality (composite score) with pricing. The most cost-efficient models achieve high benchmark scores while maintaining low per-token API costs. Free and budget-tier models that perform well are the most cost-efficient options.

Not necessarily. Our efficiency analysis shows diminishing returns beyond certain context sizes. Models with 128K tokens often score similarly to those with 1M+ tokens, meaning the extra context capacity adds cost without proportional quality gains for most use cases.

AI Model Efficiency Explorer - Score Per Context Analysis | LM Market Cap