Skip to content
基准测试类别

Reasoning 基准测试

Compare top models across the benchmark suite that best represents reasoning performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.

5

类别中的基准测试

71

有覆盖的模型

2

有人类基准的基准测试

2

饱和的基准测试

包含内容

The current benchmark set in this category, with context on what each test captures.

所有基准测试

Graduate-Level Google-Proof Q&A (Diamond)

One of the best discriminators between models. Scores range widely (40-85%), making it highly informative for comparing reasoning ability.

accuracy %人类 65%

BIG-Bench Hard

One of the best tests of structured reasoning ability. Scores range 60-95% for frontier models, providing good differentiation.

accuracy %

AI2 Reasoning Challenge (Challenge Set)

饱和

Tests commonsense scientific reasoning. Largely saturated for frontier models but still useful for comparing mid-tier and open-source models.

accuracy %

HellaSwag Commonsense NLI

饱和

Fundamental commonsense reasoning test. Saturated for frontier models (>95%) but useful for evaluating smaller models.

accuracy %人类 95.6%

Humanity's Last Exam

The hardest academic benchmark — top models still fail 60-65% of questions. Shows how far we are from genuine expert-level reasoning.

accuracy %
Model type:

Graduate-Level Google-Proof Q&A (Diamond)

Expert-level science reasoning across biology, chemistry, and physics at PhD level. Questions are designed to be 'Google-proof' — even domain experts with web access struggle.

Metric: accuracy %Human baseline: 65%

Why it matters

One of the best discriminators between models. Scores range widely (40-85%), making it highly informative for comparing reasoning ability.

GPQA Diamond Scores (57 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇Gemini 3.1 Pro Preview94.3%
2🥈Gemini 3.1 Pro94.3%
3🥉Claude Opus 4.794.2%
4Claude Fable 594.1%
5GPT-5.593.6%
6GPT-5.5 Pro93.6%
7Claude Opus 4.893.6%
8Gemini 3 Pro91.9%
9Gemini 3 Flash90.4%
10DeepSeek V4 Pro90.1%
11Grok 4.390.0%
12GPT-5.488.5%
13Claude Opus 4.688.4%
14o387.7%
15GPT-5.287.0%
16GPT-5.186.5%
17Claude Opus 4.586.2%
18DeepSeek V3.285.7%
19GPT-585.0%
20Claude 3.7 Sonnet84.8%
21Grok 384.6%
22Gemma 4 31B84.3%
23Gemini 2.5 Pro84.0%
24Claude Sonnet 4.583.4%
25Gemini 2.5 Flash82.8%
26Claude Opus 482.6%
27Claude Sonnet 4.682.0%
28Grok 482.0%
29o4-mini81.4%
30o3-mini79.7%
31o178.0%
32o1 Preview78.0%
33DeepSeek R1-052876.0%
34Claude Sonnet 475.4%
35DeepSeek R171.5%
36Llama 4 Maverick69.8%
37Claude 3.5 Sonnet65.0%
38DeepSeek V3 (March 2025)62.5%
39Gemini 2.0 Flash62.1%
40o1-mini60.0%
41Gemini 1.5 Pro59.1%
42DeepSeek V359.1%
43Llama 4 Scout57.2%
44Phi-456.1%
45GPT-4.156.0%
46GPT-4o53.6%
47Mistral Large 252.5%
48Mistral Large 252.5%
49Claude Haiku 4.552.0%
50Llama 3.1 405B51.1%
51Llama 3.3 70B50.5%
52Claude 3 Opus50.4%
53Qwen 2.5 72B49.0%
54GPT-4 Turbo48.0%
55Llama 3.1 70B46.7%
56Claude 3.5 Haiku41.6%
57GPT-4o mini40.2%

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-08T12:30:08.134Z. Some scores are approximate.

Frequently Asked Questions

AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.

Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.

A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.

Reasoning AI Benchmarks - Compare Top Models | LM Market Cap