Skip to content

AI Reasoning Benchmark

How do reasoning models stack up against standard LLMs? This benchmark compares 270 reasoning models against 116 standard models on composite score, pricing, and capabilities - helping you decide when chain-of-thought thinking is worth the trade-off.

Reasoning vs Standard - Head-to-Head

Reasoning Models
270
Models
97
Top Score
70
Avg Score
$14.12
Avg $/1M Out
Standard Models
116
Models
91
Top Score
59
Avg Score
$3.61
Avg $/1M Out

Reasoning models from 42 providers. Score difference: +11 points average for reasoning models.

Reasoning Models - Ranked by Score

#ModelScore
1Claude Fable 5Anthropic97
2Claude Fable 5 (batch)Anthropic97
3Claude Opus 5 (Fast)Anthropic95
4Claude Opus 5Anthropic95
5Claude Opus 4.8 (Fast)Anthropic95
6Claude Opus 4.8Anthropic95
7Claude Opus 4.7 (Fast)Anthropic95
8Claude Opus 4.7Anthropic95
9Claude Opus 4.7 (batch)Anthropic95
10Claude Opus 4.8 (batch)Anthropic95
11GPT-5.5 ProOpenAI93
12GPT-5.5 Pro (batch)OpenAI93
13GPT-5.5OpenAI93
14GPT-5.5 (batch)OpenAI93
15Gemini 3.1 Pro Preview Custom ToolsGoogle92
16Gemini 3.1 Pro PreviewGoogle92
17Gemini 3.1 Pro Preview (batch)Google92
18GPT-5.4 ProOpenAI92
19GPT-5.4 Pro (batch)OpenAI92
20GPT-5.4OpenAI92
21GPT-5.4 (batch)OpenAI92
22GPT-5.3-CodexOpenAI91
23GPT-5.2-CodexOpenAI91
24GPT-5.2 ProOpenAI91
25GPT-5.2 Pro (batch)OpenAI91
26GPT-5.2OpenAI91
27GPT-5.2 (batch)OpenAI91
28Claude Opus 4.6Anthropic90
29Claude Opus 4.6 (batch)Anthropic90
30GPT-5.6 Luna ProOpenAI89
31GPT-5.6 Luna Pro (batch)OpenAI89
32GPT-5.6 LunaOpenAI89
33GPT-5.6 Luna (batch)OpenAI89
34GPT-5.6 Terra ProOpenAI89
35GPT-5.6 Terra Pro (batch)OpenAI89
36GPT-5.6 TerraOpenAI89
37GPT-5.6 Terra (batch)OpenAI89
38GPT-5.6 Sol ProOpenAI89
39GPT-5.6 Sol Pro (batch)OpenAI89
40GPT-5.6 SolOpenAI89
41GPT-5.6 Sol (batch)OpenAI89
42Grok 4.5xAI89
43Grok 4.3xAI89
44Grok 4.20xAI89
45GPT-5.1-Codex-MaxOpenAI89
46GPT-5.1OpenAI89
47GPT-5.1-CodexOpenAI89
48GPT-5 ProOpenAI89
49GPT-5 Pro (batch)OpenAI89
50GPT-5 Codex (batch)OpenAI89
51GPT-5OpenAI89
52GPT-5 (batch)OpenAI89
53Gemini 3 Flash PreviewGoogle88
54Gemini 3 Flash Preview (batch)Google88
55Grok 4.20 Multi-AgentxAI88
56GPT-5.1 (batch)OpenAI88
57GPT-5.1-Codex-MiniOpenAI88
58DeepSeek V4 ProDeepSeek87
59o3 ProOpenAI87
60o3 Pro (batch)OpenAI87
61o3OpenAI87
62o3 (batch)OpenAI87
63Claude Sonnet 5Anthropic85
64Claude Sonnet 4.6Anthropic85
65Claude Sonnet 4.6 (batch)Anthropic85
66Claude Opus 4.5Anthropic85
67Claude Opus 4.5 (batch)Anthropic85
68Gemini 2.5 ProGoogle84
69Gemini 2.5 Pro (batch)Google84
70Gemini 2.5 Pro Preview 06-05Google84
71Gemini 2.5 Pro Preview 05-06Google84
72Claude Sonnet 4.5Anthropic82
73Claude Sonnet 4.5 (batch)Anthropic82
74Claude Opus 4.1Anthropic82
75Claude Opus 4Anthropic82
76o4 Mini High (batch)OpenAI81
77o4 MiniOpenAI81
78o4 Mini (batch)OpenAI81
79DeepSeek V3.2DeepSeek81
80Qwen3.8 MaxAlibaba81
81Gemma 4 31BGoogle81
82Gemma 4 31B (free)Google81
83Muse Spark 1.2meta80
84Muse Spark 1.1meta80
85Gemini 3.6 FlashGoogle80
86Gemini 3.6 Flash (batch)Google80
87Qwen3.5 397B A17BAlibaba79
88R1 0528DeepSeek79
89GPT-5.4 NanoOpenAI79
90GPT-5.4 Nano (batch)OpenAI79
91GPT-5.4 MiniOpenAI79
92GPT-5.4 Mini (batch)OpenAI79
93Gemini 3.1 Flash Lite PreviewGoogle79
94Gemini 2.5 Flash LiteGoogle79
95Gemini 2.5 Flash Lite (batch)Google79
96Gemini 2.5 FlashGoogle79
97Gemini 2.5 Flash (batch)Google79
98Gemini 3.5 FlashGoogle79
99Gemini 3.5 Flash (batch)Google79
100GLM 5.2Zhipu AI79
101GLM 5.2 (batch)Zhipu AI78
102GLM 5.1Zhipu AI78
103MiniMax M2.7MiniMax78
104GLM 5 TurboZhipu AI78
105MiniMax M2.5MiniMax78
106GLM 5Zhipu AI78
107Qwen3.5-122B-A10BAlibaba78
108Qwen3.5-27BAlibaba77
109Gemini 3.5 Flash LiteGoogle77
110Gemini 3.5 Flash Lite (batch)Google77
111GPT-5 Mini (batch)OpenAI77
112GPT-5 Nano (batch)OpenAI77
113MiMo-V2.5-ProXiaomi76
114Qwen3.5-35B-A3BAlibaba76
115Qwen3.7 PlusAlibaba76
116Kimi K2.6Moonshot AI76
117o3 MiniOpenAI75
118GLM 4.7Zhipu AI75
119GLM 4.6Zhipu AI75
120Claude Opus 4.1 (batch)Anthropic75
121GLM 4.5Zhipu AI75
122Qwen3.7 MaxAlibaba75
123Qwen3.5 Plus 2026-04-20Alibaba75
124Qwen3.6 Max PreviewAlibaba75
125Claude Sonnet 4Anthropic74
126o1OpenAI74
127o1 (batch)OpenAI74
128MiniMax M3MiniMax74
129o3 Mini High (batch)OpenAI74
130o3 Mini (batch)OpenAI74
131R1DeepSeek74
132MiniMax M3 (batch)MiniMax74
133Qwen3.6 PlusAlibaba74
134Inklingthinkingmachines74
135Hy3Tencent74
136o1-proOpenAI74
137o1-pro (batch)OpenAI74
138MiMo-V2.5Xiaomi73
139Gemma 4 26B A4B Google73
140Gemma 4 26B A4B (free)Google73
141Inkling (batch)thinkingmachines73
142GLM 5V TurboZhipu AI72
143o4 Mini HighOpenAI72
144MiniMax M2.1MiniMax72
145MiniMax M2MiniMax72
146DeepSeek V3.2 ExpDeepSeek72
147DeepSeek V3.1DeepSeek72
148Mistral Medium 3.5Mistral AI72
149Inkling Smallthinkingmachines72
150MiniMax M1MiniMax71
151GLM 4.5 AirZhipu AI71
152Claude Haiku 4.5Anthropic70
153Claude Haiku 4.5 (batch)Anthropic70
154Qwen3 VL 235B A22B ThinkingAlibaba69
155DeepSeek V3.1 TerminusDeepSeek69
156Qwen3.5-FlashAlibaba69
157Hy3 previewTencent68
158Qwen3.5 Plus 2026-02-15Alibaba68
159Qwen3 Max ThinkingAlibaba68
160Qwen3 Next 80B A3B ThinkingAlibaba67
161Qwen3.5-9BAlibaba67
162Step 3.5 FlashStepFun66
163Composer 2Cursor66
164Composer 2 FastCursor66
165GLM 4.6VZhipu AI66
166Qwen3 235B A22B Thinking 2507Alibaba66
167Qwen3 30B A3B Thinking 2507Alibaba64
168Qwen3 30B A3BAlibaba64
169o3 Mini HighOpenAI64
170Trinity Large Thinkingarcee-ai64
171GPT-5 MiniOpenAI64
172GLM 4.7 FlashZhipu AI64
173GLM 4.5VZhipu AI63
174Mercury 2Inception61
175Qwen3 8BAlibaba61
176Nova 2 LiteAmazon61
177Kimi K2.5Moonshot AI59
178gpt-oss-20bOpenAI57
179gpt-oss-20b (free)OpenAI57
180Olmo 3 32B ThinkAllen AI55
181Kimi K2.7 CodeMoonshot AI54
182Kimi K2.7 Code (batch)Moonshot AI54
183Qwen3 235B A22BAlibaba54
184Kimi K2 ThinkingMoonshot AI53
185GPT-5 NanoOpenAI47
186R1 Distill Llama 70BDeepSeek41
187gpt-oss-120bOpenAI40
188Ling 3.0 Tiny (free)inclusionai40
189DeepSeek V4 Flash Latest~deepseek40
190DeepSeek V4 Flash 0731DeepSeek40
191Qwen3.7 FlashAlibaba40
192Claude Opus 5 (batch)Anthropic40
193Ling-3.0-flashinclusionai40
194Laguna S 2.1poolside40
195Laguna S 2.1 (free)poolside40
196LongCat 2.0Meituan40
197Kimi K3Moonshot AI40
198Grok Latest~x-ai40
199Aion-3.0-Miniaion-labs40
200Aion-3.0aion-labs40
201Laguna XS 2.1poolside40
202Laguna XS 2.1 (free)poolside40
203Claude Sonnet 5 (batch)Anthropic40
204Fugu Ultrasakana40
205North Mini Code (free)Cohere40
206Claude Fable Latest~anthropic40
207Nemotron 3.5 Content Safety (free)NVIDIA40
208Nemotron 3 UltraNVIDIA40
209Nemotron 3 Ultra (batch)NVIDIA40
210Nemotron 3 Ultra (free)NVIDIA40
211Step 3.7 FlashStepFun40
212Grok Build 0.1xAI40
213Perceptron Mk1perceptron40
214Ring-2.6-1Tinclusionai40
215Nemotron 3 Nano Omni (free)NVIDIA40
216Anthropic Claude Haiku Latest~anthropic40
217OpenAI GPT Mini Latest~openai40
218Google Gemini Pro Latest~google40
219MoonshotAI Kimi Latest~moonshotai40
220Google Gemini Flash Latest~google40
221Anthropic Claude Sonnet Latest~anthropic40
222OpenAI GPT Latest~openai40
223Qwen3.6 FlashAlibaba40
224Qwen3.6 35B A3BAlibaba40
225Qwen3.6 27BAlibaba40
226DeepSeek V4 Flash 0423DeepSeek40
227Claude Opus Latest~anthropic40
228Mistral Small 4Mistral AI40
229Nemotron 3 SuperNVIDIA40
230Nemotron 3 Super (free)NVIDIA40
231Seed-2.0-LiteByteDance40
232Seed-2.0-MiniByteDance40
233Aion-2.0aion-labs40
234Solar Pro 3Upstage40
235Seed 1.6 FlashByteDance40
236Seed 1.6ByteDance40
237Nemotron 3 Nano 30B A3BNVIDIA40
238Nemotron 3 Nano 30B A3B (free)NVIDIA40
239Falcon-H1-Arabic 34B InstructTII40
240Falcon-H1-Arabic 7B InstructTII40
241Falcon-H1-Arabic 3B InstructTII40
242Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google-
243Nano Banana 2 (Gemini 3.1 Flash Image)Google-
244Nano Banana Pro (Gemini 3 Pro Image)Google-
245Gemini 3.1 Flash LiteGoogle-
246Gemini 3.1 Flash Lite (batch)Google-
247GPT-5.4 Image 2OpenAI-
248Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google-
249Nano Banana Pro (Gemini 3 Pro Image Preview)Google-
250Cogito v2.1 671Bdeepcogito-
251Sonar Pro SearchPerplexity-
252gpt-oss-safeguard-20bOpenAI-
253Nemotron Nano 12B 2 VL (free)NVIDIA-
254GPT-5 Image MiniOpenAI-
255Qwen3 VL 8B ThinkingAlibaba-
256GPT-5 ImageOpenAI-
257Qwen3 VL 30B A3B ThinkingAlibaba-
258Qwen Plus 0728 (thinking)Alibaba-
259Nemotron Nano 9B V2 (free)NVIDIA-
260Hunyuan A13B InstructTencent-
261ERNIE 4.5 VL 424B A47B Baidu-
262Qwen3 14BAlibaba-
263Qwen3 32BAlibaba-
264Reka Flash 3rekaai-
265Sonar Reasoning ProPerplexity-
266Sonar Deep ResearchPerplexity-
267SWE-1.5Windsurf-
268Falcon Arabic 7B InstructTII-
269Falcon3 10B InstructTII-
270Falcon3 7B InstructTII-

Top Standard Models (Non-Reasoning) - For Comparison

#ModelScore
1GPT-5.3 ChatOpenAI91
2GPT-5.2 ChatOpenAI91
3Gemma 2 27BGoogle77
4GPT-4o (batch)OpenAI72
5DeepSeek V3 0324DeepSeek72
6GPT-4o (2024-11-20)OpenAI71
7GPT-4o (2024-08-06)OpenAI71
8GPT-4oOpenAI71
9GPT-4o (2024-05-13)OpenAI71
10DeepSeek V3DeepSeek70

Understanding AI Reasoning Benchmarks

What Is Chain-of-Thought?

Chain-of-thought (CoT) prompting enables AI models to break down complex problems into intermediate steps before producing a final answer. Models like OpenAI o1 and DeepSeek R1 internalize this process, generating hidden reasoning traces that dramatically improve accuracy on math, logic, and multi-step tasks compared to direct answering.

When Reasoning Helps

Reasoning models shine on tasks that require multiple logical steps: mathematical proofs, complex coding challenges, scientific analysis, strategic planning, and any problem where standard models tend to hallucinate or skip steps. For simple Q&A or creative writing, standard models are often faster and equally effective.

Speed vs Accuracy

Reasoning models consume more tokens and take longer to respond because they generate internal thinking traces. This trade-off is worthwhile when correctness matters more than latency - for example in code generation, financial analysis, or exam-style problems. For real-time chat, standard models remain the better choice.

Emerging Reasoning Models

The reasoning model landscape is evolving rapidly. OpenAI's o1 and o3 series led the way, followed by DeepSeek R1 bringing open-source reasoning. Google, Anthropic, and other providers have since introduced their own reasoning-capable models, driving down costs and expanding access to chain-of-thought capabilities.

Frequently Asked Questions

AI reasoning benchmarks test a model's ability to solve complex problems requiring logical thinking, mathematical reasoning, scientific analysis, and multi-step problem solving - tasks that go beyond simple pattern matching.

DeepSeek R1, OpenAI o3, and Claude with extended thinking lead on reasoning benchmarks. These models use chain-of-thought processing to break down complex problems into steps, achieving significantly higher accuracy.

Key reasoning benchmarks include GPQA Diamond (graduate-level science), MATH-500 (mathematical reasoning), AIME (competition math), ARC Challenge (science questions), and GSM8K (grade-school math). Each tests different aspects of reasoning ability.

AI Reasoning Benchmark - Which LLM Thinks Best? | LM Market Cap