Skip to content
基准测试类别

Knowledge 基准测试

Compare top models across the benchmark suite that best represents knowledge performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.

3

类别中的基准测试

88

有覆盖的模型

1

有人类基准的基准测试

1

饱和的基准测试

包含内容

The current benchmark set in this category, with context on what each test captures.

所有基准测试

Massive Multitask Language Understanding

饱和

Shows how well a model has absorbed factual knowledge during training. Saturating above 90%, so less useful for differentiating frontier models.

accuracy %人类 89.8%

MMLU Professional

Better at differentiating top models since scores are 16-33% lower than standard MMLU. Tests reasoning in addition to knowledge.

accuracy %

SimpleQA Factual Accuracy

GPT-4o scores below 40%, making it surprisingly challenging. Tests honesty and factual reliability, not just knowledge breadth.

accuracy %
Model type:

Massive Multitask Language Understanding

Tests broad knowledge across 57 academic subjects (STEM, humanities, social sciences) with 16,000 multiple-choice questions. The most widely-cited LLM benchmark.

Metric: accuracy %Human baseline: 89.8%Saturated

Why it matters

Shows how well a model has absorbed factual knowledge during training. Saturating above 90%, so less useful for differentiating frontier models.

MMLU Scores (54 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇GPT-5.494.0%
2🥈GPT-5.293.5%
3🥉GPT-5.193.2%
4GPT-593.0%
5Gemini 3.1 Pro Preview92.6%
6Gemini 3.1 Pro92.6%
7Gemini 3 Pro92.5%
8GPT-5.592.4%
9GPT-5.5 Pro92.4%
10o392.3%
11Claude Opus 4.692.1%
12o191.8%
13DeepSeek R1-052891.5%
14Grok 491.5%
15Claude Opus 4.591.4%
16Claude Sonnet 4.691.2%
17Claude Opus 491.0%
18Gemini 2.5 Pro90.8%
19DeepSeek R190.8%
20Claude Sonnet 4.590.8%
21o1 Preview90.8%
22Claude 3.7 Sonnet90.2%
23Claude Sonnet 489.5%
24GPT-4.189.2%
25DeepSeek V3 (March 2025)89.2%
26GPT-4o88.7%
27Claude 3.5 Sonnet88.7%
28Llama 3.1 405B88.6%
29DeepSeek V388.5%
30Grok 388.5%
31DeepSeek V3.288.5%
32Llama 4 Maverick88.0%
33Gemini 3 Flash88.0%
34Grok 287.5%
35o3-mini86.9%
36Claude 3 Opus86.8%
37GPT-4 Turbo86.5%
38Llama 3.3 70B86.3%
39Qwen 2.5 72B86.1%
40Llama 3.1 70B86.0%
41Gemini 1.5 Pro85.9%
42Gemini 2.5 Flash85.8%
43o1-mini85.2%
44Phi-484.8%
45Mistral Large 284.7%
46Claude Haiku 4.584.5%
47Mistral Large 284.0%
48GPT-4o mini82.0%
49Claude 3.5 Haiku80.9%
50Llama 4 Scout79.6%
51Mixtral 8x22B77.3%
52Gemini 2.0 Flash76.4%
53Command R+75.7%
54Gemma 2 27B75.2%

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-08T12:30:08.134Z. Some scores are approximate.

Frequently Asked Questions

AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.

Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.

A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.

Knowledge AI Benchmarks - Compare Top Models | LM Market Cap