Skip to content

上下文窗口效率探索器

分析300个AI模型的评分/上下文令牌比率,找出最充分利用上下文窗口、输出容量和成本的模型。

上下文窗口 vs 评分

LMMarketCap.com

效率概览

所有分析模型的关键效率指标。

最高效 (128K+)

GPT-5.2 Chat

707.0 score/MToken

最佳输出效率

Gemma 2 27B

37.8 score/1K output

最佳成本效率

gpt-oss-20b

1063.0 score/$

平均总体效率

6.0%

所有模型标准化

效率排名

按每百万上下文令牌评分排名的前50个模型。

#模型评分上下文评分/百万令牌
1Gemma 2 27BGoogle778K9448.2
2GPT-4OpenAI658K7923.3
3R1 Distill Llama 70BDeepSeek418K4968.3
4Hy-MT2-1.8BTencent408K4882.8
5Hy-MT2-30B-A3BTencent408K4882.8
6Hy-MT2-7BTencent408K4882.8
7Phi 4Microsoft6016K3674.3
8GLM 5.2 (free)Zhipu AI7633K2310.2
9R1DeepSeek7464K1153.1
10MiniMax M2-herMiniMax6866K1039.1
11Mixtral 8x22B InstructMistral AI6366K967.4
12GLM 4.5VZhipu AI6266K950.6
13Qwen3 30B A3B Thinking 2507Alibaba6482K782.5
14GPT-5.2 ChatOpenAI91128K707.0
15LFM2.5-2.6B (free)Liquid AI4066K610.4
16GLM 4.5Zhipu AI75131K573.0
17GPT-4o (batch)OpenAI72128K562.5
18GPT-4o (2024-11-20)OpenAI71128K556.3
19GPT-4o (2024-08-06)OpenAI71128K556.3
20GPT-4oOpenAI71128K556.3
21GPT-4o (2024-05-13)OpenAI71128K556.3
22GPT-4o-miniOpenAI69128K541.4
23GLM 4.5 AirZhipu AI71131K539.4
24Qwen3 VL 235B A22B ThinkingAlibaba69131K528.7
25GPT-4 TurboOpenAI67128K521.1
26Mistral LargeMistral AI66128K514.8
27Llama 3.3 70B InstructMeta67131K509.6
28DeepSeek V3.2DeepSeek83164K509.0
29Mistral Large 2407Mistral AI66131K502.8
30GLM 4.6VZhipu AI66131K500.5
31Qwen3 235B A22B Thinking 2507Alibaba65131K498.2
32Llama 3.1 70B InstructMeta65131K498.2
33Qwen3 30B A3BAlibaba64131K489.0
34GPT-4o-mini (batch)OpenAI63128K488.3
35R1 0528DeepSeek79164K484.6
36GPT-4 Turbo (batch)OpenAI62128K480.5
37Mercury 2Inception61128K475.8
38gpt-oss-120bOpenAI62131K470.0
39Qwen3 8BAlibaba61131K465.4
40GPT-4o-mini (2024-07-18)OpenAI57128K441.4
41DeepSeek V3.2 ExpDeepSeek72164K438.2
42DeepSeek V3.1DeepSeek72164K438.2
43DeepSeek V3 0324DeepSeek72164K438.2
44gpt-oss-20bOpenAI57131K437.9
45gpt-oss-20b (batch)OpenAI57131K437.9
46o3 ProOpenAI87200K433.5
47o3OpenAI87200K433.5
48o3 (batch)OpenAI87200K433.5
49Claude Opus 4.5Anthropic85200K425.5
50Claude Opus 4.5 (batch)Anthropic85200K425.5

层级分析

不同上下文窗口层级的效率分析。

Small7 models
平均评分52
评分/百万令牌5808.9

最差

Phi 4

Medium7 models
平均评分64
评分/百万令牌1116.2
Large146 models
平均评分68
评分/百万令牌321.8
Mega140 models
平均评分70
评分/百万令牌66.7

边际递减分析

更大的上下文窗口是否与更高的评分相关?

层级平均上下文平均评分平均效率
Small9K525808.9
Medium63K641116.2
Large240K68321.8
Mega1.1M7066.7

输出令牌效率

按输出效率(每1K输出令牌评分)排名的前20个模型。16K+输出令牌的模型已高亮显示。

模型评分最大输出输出效率
Gemma 2 27BGoogle772K37.8
MiniMax M2-herMiniMax682K33.3
GPT-4o (2024-05-13)OpenAI714K17.4
GPT-4 TurboOpenAI674K16.3
GPT-4OpenAI654K15.8
GPT-4 Turbo (batch)OpenAI624K15.0
Claude 3 HaikuAnthropic514K12.5
Command R (08-2024)Cohere494K12.2
Command R+ (08-2024)Cohere494K12.2
Schematron V2 Smallinference-net404K9.8
Hy-MT2-1.8BTencent404K9.8
Hy-MT2-30B-A3BTencent404K9.8
Hy-MT2-7BTencent404K9.8
Qwen3 8BAlibaba618K7.4
Qwen3 235B A22BAlibaba548K6.6
Command ACohere518K6.2
R1 Distill Llama 70BDeepSeek417K5.5
Gemma 4 31BGoogle16K+8116K4.9
Schematron V2 Turboinference-net408K4.9
LFM2.5-2.6B (free)Liquid AI408K4.9

关键洞察

从效率数据中自动生成的观察结果。

上下文最优点

Small models have the highest average efficiency at 5808.9 score/MToken across 7 models.

输出很重要

Models with 16K+ output tokens score 30% higher on average than models with smaller output limits.

紧凑型高性能模型

0 models achieve top-20 scores with under 128K context.

探索更多

深入了解上下文窗口、对比模型或探索其他维度。

Frequently Asked Questions

Efficiency is measured as the score-per-context-token ratio - how much ranking score a model achieves relative to its context window size. Models that score highly with smaller context windows are considered more efficient than those requiring massive context to achieve similar results.

Cost efficiency combines quality (composite score) with pricing. The most cost-efficient models achieve high benchmark scores while maintaining low per-token API costs. Free and budget-tier models that perform well are the most cost-efficient options.

Not necessarily. Our efficiency analysis shows diminishing returns beyond certain context sizes. Models with 128K tokens often score similarly to those with 1M+ tokens, meaning the extra context capacity adds cost without proportional quality gains for most use cases.

AI Model Efficiency Explorer - Score Per Context Analysis | LM Market Cap