Skip to content

AI Benchmarks - LLM Performance Scores

Last updated: 24m ago

Real benchmark scores from official model cards and third-party evaluations. Compare 166 models across 21 benchmarks - from MMLU and GPQA Diamond to SWE-bench and Arena Elo. Filter by category, model type, and switch between chart and matrix views.

Benchmark Categories

Jump directly into the strongest benchmark clusters instead of starting from the full matrix.

Model type:

Massive Multitask Language Understanding

Tests broad knowledge across 57 academic subjects (STEM, humanities, social sciences) with 16,000 multiple-choice questions. The most widely-cited LLM benchmark.

Metric: accuracy %Human baseline: 89.8%Saturated

Why it matters

Shows how well a model has absorbed factual knowledge during training. Saturating above 90%, so less useful for differentiating frontier models.

MMLU Scores (54 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇GPT-5.494.0%
2🥈GPT-5.293.5%
3🥉GPT-5.193.2%
4GPT-593.0%
5Gemini 3.1 Pro Preview92.6%
6Gemini 3.1 Pro92.6%
7Gemini 3 Pro92.5%
8GPT-5.592.4%
9GPT-5.5 Pro92.4%
10o392.3%
11Claude Opus 4.692.1%
12o191.8%
13DeepSeek R1-052891.5%
14Grok 491.5%
15Claude Opus 4.591.4%
16Claude Sonnet 4.691.2%
17Claude Opus 491.0%
18Gemini 2.5 Pro90.8%
19DeepSeek R190.8%
20Claude Sonnet 4.590.8%
21o1 Preview90.8%
22Claude 3.7 Sonnet90.2%
23Claude Sonnet 489.5%
24GPT-4.189.2%
25DeepSeek V3 (March 2025)89.2%
26GPT-4o88.7%
27Claude 3.5 Sonnet88.7%
28Llama 3.1 405B88.6%
29DeepSeek V388.5%
30Grok 388.5%
31DeepSeek V3.288.5%
32Llama 4 Maverick88.0%
33Gemini 3 Flash88.0%
34Grok 287.5%
35o3-mini86.9%
36Claude 3 Opus86.8%
37GPT-4 Turbo86.5%
38Llama 3.3 70B86.3%
39Qwen 2.5 72B86.1%
40Llama 3.1 70B86.0%
41Gemini 1.5 Pro85.9%
42Gemini 2.5 Flash85.8%
43o1-mini85.2%
44Phi-484.8%
45Mistral Large 284.7%
46Claude Haiku 4.584.5%
47Mistral Large 284.0%
48GPT-4o mini82.0%
49Claude 3.5 Haiku80.9%
50Llama 4 Scout79.6%
51Mixtral 8x22B77.3%
52Gemini 2.0 Flash76.4%
53Command R+75.7%
54Gemma 2 27B75.2%

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-09T12:30:07.628Z. Some scores are approximate.

Beyond Benchmarks

Benchmarks tell part of the story. Our composite score combines real capabilities, pricing, context window, and more. Compare models head-to-head or explore the full leaderboard for a complete picture.

Frequently Asked Questions

AI benchmarks are standardized tests that measure how well AI models perform at specific tasks. Common benchmarks include MMLU (general knowledge), SWE-bench (coding), GPQA (science reasoning), MATH-500 (math), Arena Elo (human preference), and HumanEval (code generation).

No single benchmark captures the full picture. MMLU tests breadth of knowledge, SWE-bench tests real-world coding ability, and Arena Elo reflects human preferences. We recommend looking at multiple benchmarks together, which is why our composite score weighs several dimensions.

Benchmark refresh cadence varies by source. We show the latest source-backed timestamp on the page so you can see when the benchmark feed was last refreshed.

Benchmarks are useful indicators but not perfect predictors. A model scoring well on MMLU may not be the best for creative writing, and high SWE-bench scores do not guarantee faster coding assistance. Real-world performance depends on your specific use case, prompt engineering, and integration approach.

AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap