Skip to content

Best AI Models for Reasoning

AI models ranked by reasoning ability using GPQA, ARC-Challenge, BIG-Bench Hard, and Humanity's Last Exam scores.

Last updated: 11m ago
#1 Model

GPT-4o

Score: 76.5

Average Score

46.8

Across all ranked models

Models Ranked

45

With benchmark data

Weights:GPQA (40%)ARC-Challenge (20%)BIG-Bench Hard (20%)Humanity's Last Exam (20%)

Top Best for Reasoning Models by Weighted Score

Top 15 models by weighted score

LMMarketCap.com

Benchmark Breakdown

Per-benchmark scores for top 10 models

GPQA
ARC-Challenge
BIG-Bench Hard
Humanity's Last Exam
LMMarketCap.com
#ModelScore
1GPT-4oOpenAI76.5
2GPT-4o-miniOpenAI75.1
3Llama 3.1 70B InstructMeta74.8
4Gemma 2 27BGoogle72.2
5DeepSeek V3 0324DeepSeek68.2
6DeepSeek V3DeepSeek67.8
7R1 0528DeepSeek67
8R1DeepSeek65.9
9GPT-4 TurboOpenAI64.4
10Llama 3.3 70B InstructMeta64.2
11Mistral LargeMistral AI62
12Claude Haiku 4.5Anthropic60.8
13Llama 4 ScoutMeta58.9
14GPT-5.4OpenAI55.7
15Claude Opus 4.6Anthropic55.1
16GPT-5.2OpenAI54.6
17GPT-5OpenAI53.3
18Gemini 2.5 ProGoogle52.4
19o3OpenAI52.3
20Gemini 3 Flash PreviewGoogle52.1
21Claude Opus 4.5Anthropic51.9
22Claude Sonnet 4.6Anthropic51.1
23Claude Opus 4Anthropic49.9
24Phi 4Microsoft49.7
25GPT-5.1OpenAI48.7
26o3 MiniOpenAI46.2
27o4 MiniOpenAI46.1
28Claude Fable 5Anthropic45.7
29Claude Sonnet 4.5Anthropic43.4
30Gemini 2.5 FlashGoogle41.3
31o1OpenAI41.3
32Gemma 4 31BGoogle39.9
33Claude Sonnet 4Anthropic39.3
34Claude Opus 4.8Anthropic38.6
35Llama 4 MaverickMeta38.3
36GPT-4.1OpenAI38
37Claude Opus 4.7Anthropic36.3
38GPT-5.5OpenAI32.1
39GPT-5 MiniOpenAI15.1
40Command R7B (12-2024)Cohere14.6
41Qwen2.5 7B InstructAlibaba13
42Llama 3.1 8B InstructMeta12.9
43Llama 3.2 3B InstructMeta10.3
44Gemini 3.1 Flash LiteGoogle6.7
45Mistral Medium 3Mistral AI3.5

How scores are calculated

Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.

Frequently Asked Questions

Based on our benchmark analysis, GPT-4o by OpenAI is currently the #1 ranked model for reasoning, with a weighted score of 76.5/100.

Models are ranked using a weighted average of GPQA, ARC-Challenge, BIG-Bench Hard, Humanity's Last Exam benchmark scores. All scores are normalized to a 0-100 scale.

We currently rank 45 models that have relevant benchmark data for reasoning tasks.

Best AI Models for Reasoning (2026) | LM Market Cap