Skip to content

AI基准测试 - LLM性能评分

Last updated: 5h ago

来自官方模型卡片和第三方评估的真实基准测试分数。比较166个模型在21个基准测试中的表现--从MMLU和GPQA Diamond到SWE-bench和Arena Elo。按类别、模型类型筛选,在图表和矩阵视图之间切换。

基准测试类别

直接跳转到最强的基准测试集群,而不是从完整矩阵开始。

Model type:

Massive Multitask Language Understanding

Tests broad knowledge across 57 academic subjects (STEM, humanities, social sciences) with 16,000 multiple-choice questions. The most widely-cited LLM benchmark.

Metric: accuracy %Human baseline: 89.8%Saturated

Why it matters

Shows how well a model has absorbed factual knowledge during training. Saturating above 90%, so less useful for differentiating frontier models.

MMLU Scores (54 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇GPT-5.494.0%
2🥈GPT-5.293.5%
3🥉GPT-5.193.2%
4GPT-593.0%
5Gemini 3.1 Pro Preview92.6%
6Gemini 3.1 Pro92.6%
7Gemini 3 Pro92.5%
8GPT-5.592.4%
9GPT-5.5 Pro92.4%
10o392.3%
11Claude Opus 4.692.1%
12o191.8%
13DeepSeek R1-052891.5%
14Grok 491.5%
15Claude Opus 4.591.4%
16Claude Sonnet 4.691.2%
17Claude Opus 491.0%
18Gemini 2.5 Pro90.8%
19DeepSeek R190.8%
20Claude Sonnet 4.590.8%
21o1 Preview90.8%
22Claude 3.7 Sonnet90.2%
23Claude Sonnet 489.5%
24GPT-4.189.2%
25DeepSeek V3 (March 2025)89.2%
26GPT-4o88.7%
27Claude 3.5 Sonnet88.7%
28Llama 3.1 405B88.6%
29DeepSeek V388.5%
30Grok 388.5%
31DeepSeek V3.288.5%
32Llama 4 Maverick88.0%
33Gemini 3 Flash88.0%
34Grok 287.5%
35o3-mini86.9%
36Claude 3 Opus86.8%
37GPT-4 Turbo86.5%
38Llama 3.3 70B86.3%
39Qwen 2.5 72B86.1%
40Llama 3.1 70B86.0%
41Gemini 1.5 Pro85.9%
42Gemini 2.5 Flash85.8%
43o1-mini85.2%
44Phi-484.8%
45Mistral Large 284.7%
46Claude Haiku 4.584.5%
47Mistral Large 284.0%
48GPT-4o mini82.0%
49Claude 3.5 Haiku80.9%
50Llama 4 Scout79.6%
51Mixtral 8x22B77.3%
52Gemini 2.0 Flash76.4%
53Command R+75.7%
54Gemma 2 27B75.2%

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-10T18:30:07.477Z. Some scores are approximate.

超越基准测试

基准测试只讲述了部分故事。我们的综合评分结合了真实功能、定价、上下文窗口等。进行模型对比或浏览完整排行榜以获得全面了解。

Frequently Asked Questions

AI基准测试是衡量AI模型在特定任务上表现的标准化测试。常见基准测试包括MMLU(通用知识)、SWE-bench(编程)、GPQA(科学推理)、MATH-500(数学)、Arena Elo(人类偏好)和HumanEval(代码生成)。

没有单一基准测试能捕捉全貌。MMLU测试知识广度,SWE-bench测试现实世界编程能力,Arena Elo反映人类偏好。我们建议综合查看多个基准测试,这就是为什么我们的综合评分权衡了多个维度。

我们的基准测试数据每小时从服务商API和社区评估中刷新。新的基准测试会在成为行业标准后添加。Arena Elo评分根据用户投票持续更新。

基准测试是有用的指标但不是完美的预测器。在MMLU上得分高的模型不一定最适合创意写作,SWE-bench高分也不保证更快的编程辅助。现实世界表现取决于您的具体使用场景、提示工程和集成方法。

AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap