Best AI Models for Reasoning
AI models ranked by reasoning ability using GPQA, ARC-Challenge, BIG-Bench Hard, and Humanity's Last Exam scores.
GPT-4o
Score: 76.5
46.8
Across all ranked models
45
With benchmark data
Top Best for Reasoning Models by Weighted Score
Top 15 models by weighted score
Benchmark Breakdown
Per-benchmark scores for top 10 models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4oOpenAI | 76.5 |
| 2 | GPT-4o-miniOpenAI | 75.1 |
| 3 | Llama 3.1 70B InstructMeta | 74.8 |
| 4 | Gemma 2 27BGoogle | 72.2 |
| 5 | DeepSeek V3 0324DeepSeek | 68.2 |
| 6 | DeepSeek V3DeepSeek | 67.8 |
| 7 | R1 0528DeepSeek | 67 |
| 8 | R1DeepSeek | 65.9 |
| 9 | GPT-4 TurboOpenAI | 64.4 |
| 10 | Llama 3.3 70B InstructMeta | 64.2 |
| 11 | Mistral LargeMistral AI | 62 |
| 12 | Claude Haiku 4.5Anthropic | 60.8 |
| 13 | Llama 4 ScoutMeta | 58.9 |
| 14 | GPT-5.4OpenAI | 55.7 |
| 15 | Claude Opus 4.6Anthropic | 55.1 |
| 16 | GPT-5.2OpenAI | 54.6 |
| 17 | GPT-5OpenAI | 53.3 |
| 18 | Gemini 2.5 ProGoogle | 52.4 |
| 19 | o3OpenAI | 52.3 |
| 20 | Gemini 3 Flash PreviewGoogle | 52.1 |
| 21 | Claude Opus 4.5Anthropic | 51.9 |
| 22 | Claude Sonnet 4.6Anthropic | 51.1 |
| 23 | Claude Opus 4Anthropic | 49.9 |
| 24 | Phi 4Microsoft | 49.7 |
| 25 | GPT-5.1OpenAI | 48.7 |
| 26 | o3 MiniOpenAI | 46.2 |
| 27 | o4 MiniOpenAI | 46.1 |
| 28 | Claude Fable 5Anthropic | 45.7 |
| 29 | Claude Sonnet 4.5Anthropic | 43.4 |
| 30 | Gemini 2.5 FlashGoogle | 41.3 |
| 31 | o1OpenAI | 41.3 |
| 32 | Gemma 4 31BGoogle | 39.9 |
| 33 | Claude Sonnet 4Anthropic | 39.3 |
| 34 | Claude Opus 4.8Anthropic | 38.6 |
| 35 | Llama 4 MaverickMeta | 38.3 |
| 36 | GPT-4.1OpenAI | 38 |
| 37 | Claude Opus 4.7Anthropic | 36.3 |
| 38 | GPT-5.5OpenAI | 32.1 |
| 39 | GPT-5 MiniOpenAI | 15.1 |
| 40 | Command R7B (12-2024)Cohere | 14.6 |
| 41 | Qwen2.5 7B InstructAlibaba | 13 |
| 42 | Llama 3.1 8B InstructMeta | 12.9 |
| 43 | Llama 3.2 3B InstructMeta | 10.3 |
| 44 | Gemini 3.1 Flash LiteGoogle | 6.7 |
| 45 | Mistral Medium 3Mistral AI | 3.5 |
How scores are calculated
Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.
Other Specialty Leaderboards
Based on our benchmark analysis, GPT-4o by OpenAI is currently the #1 ranked model for reasoning, with a weighted score of 76.5/100.
Models are ranked using a weighted average of GPQA, ARC-Challenge, BIG-Bench Hard, Humanity's Last Exam benchmark scores. All scores are normalized to a 0-100 scale.
We currently rank 45 models that have relevant benchmark data for reasoning tasks.