Skip to content

AI Reasoning Benchmark

How do reasoning models stack up against standard LLMs? This benchmark compares 322 reasoning models against 116 standard models on composite score, pricing, and capabilities - helping you decide when chain-of-thought thinking is worth the trade-off.

Reasoning vs Standard - Head-to-Head

Reasoning Models
322
Models
96
Top Score
70
Avg Score
$10.80
Avg $/1M Out
Standard Models
116
Models
91
Top Score
61
Avg Score
$3.23
Avg $/1M Out

Reasoning models from 45 providers. Score difference: +9 points average for reasoning models.

Reasoning Models - Ranked by Score

#ModelScore
1Claude Fable 5.1Anthropic96
2Claude Fable 5.1 (batch)Anthropic96
3Claude Fable 5Anthropic96
4Claude Fable 5 (batch)Anthropic96
5Claude Opus 5.5Anthropic95
6Claude Opus 5Anthropic95
7Claude Opus 4.8Anthropic95
8Claude Opus 4.7Anthropic95
9Claude Opus 4.7 (batch)Anthropic95
10Claude Opus 4.8 (batch)Anthropic95
11GPT-5.5 ProOpenAI93
12GPT-5.5 Pro (batch)OpenAI93
13GPT-5.5OpenAI93
14GPT-5.5 (batch)OpenAI93
15Gemini 3.1 Pro Preview Custom ToolsGoogle92
16Gemini 3.1 Pro PreviewGoogle92
17Gemini 3.1 Pro Preview (batch)Google92
18GPT-5.4 ProOpenAI92
19GPT-5.4 Pro (batch)OpenAI92
20GPT-5.4OpenAI92
21GPT-5.4 (batch)OpenAI92
22GPT-5.3-CodexOpenAI91
23GPT-5.2-CodexOpenAI91
24GPT-5.2 ProOpenAI91
25GPT-5.2 Pro (batch)OpenAI91
26GPT-5.2OpenAI91
27GPT-5.2 (batch)OpenAI91
28Claude Opus 4.6Anthropic90
29Claude Opus 4.6 (batch)Anthropic90
30GPT-5.6 Luna ProOpenAI89
31GPT-5.6 Luna Pro (batch)OpenAI89
32GPT-5.6 LunaOpenAI89
33GPT-5.6 Luna (batch)OpenAI89
34GPT-5.6 Terra ProOpenAI89
35GPT-5.6 Terra Pro (batch)OpenAI89
36GPT-5.6 TerraOpenAI89
37GPT-5.6 Terra (batch)OpenAI89
38GPT-5.6 Sol ProOpenAI89
39GPT-5.6 Sol Pro (batch)OpenAI89
40GPT-5.6 SolOpenAI89
41GPT-5.6 Sol (batch)OpenAI89
42Grok 4.7xAI89
43Grok 4.6xAI89
44Grok 4.5xAI89
45Grok 4.3xAI89
46Grok 4.20xAI89
47GPT-5.1-Codex-MaxOpenAI89
48GPT-5.1OpenAI89
49GPT-5.1-CodexOpenAI89
50GPT-5 ProOpenAI89
51GPT-5 Pro (batch)OpenAI89
52GPT-5OpenAI89
53GPT-5 (batch)OpenAI89
54Gemini 3 Flash PreviewGoogle88
55Gemini 3 Flash Preview (batch)Google88
56Grok 4.20 Multi-AgentxAI88
57GPT-5.1 (batch)OpenAI88
58GPT-5.1-Codex-MiniOpenAI88
59o3 ProOpenAI87
60o3OpenAI87
61o3 (batch)OpenAI87
62Claude Sonnet 5Anthropic85
63Claude Sonnet 4.6Anthropic85
64Claude Sonnet 4.6 (batch)Anthropic85
65Claude Opus 4.5Anthropic85
66Claude Opus 4.5 (batch)Anthropic85
67Gemini 2.5 ProGoogle84
68Gemini 2.5 Pro (batch)Google84
69Gemini 2.5 Pro Preview 06-05Google84
70DeepSeek V3.2DeepSeek83
71Claude Sonnet 4.5Anthropic82
72Claude Sonnet 4.5 (batch)Anthropic82
73GPT-6 AstraOpenAI82
74GPT-6 Astra (batch)OpenAI82
75GPT-6 Astra ProOpenAI82
76GPT-6 Astra Pro (batch)OpenAI82
77o4 MiniOpenAI81
78o4 Mini (batch)OpenAI81
79Gemini 3.8 FlashGoogle81
80Gemini 3.8 Flash (batch)Google81
81Grok 4.3 (batch)xAI81
82Muse Spark 1.3meta81
83Muse Spark 1.2meta81
84Muse Spark 1.1meta81
85Gemma 4 31BGoogle81
86Gemma 4 31B (free)Google81
87Qwen3.5 397B A17BAlibaba79
88R1 0528DeepSeek79
89GPT-5.4 NanoOpenAI79
90GPT-5.4 Nano (batch)OpenAI79
91GPT-5.4 MiniOpenAI79
92GPT-5.4 Mini (batch)OpenAI79
93Gemini 3.1 Flash Lite PreviewGoogle79
94Gemini 2.5 Flash LiteGoogle79
95Gemini 2.5 Flash Lite (batch)Google79
96Gemini 2.5 FlashGoogle79
97Gemini 2.5 Flash (batch)Google79
98Gemini 3.5 FlashGoogle79
99Gemini 3.5 Flash (batch)Google79
100GLM 5.3Zhipu AI79
101GLM 5.3 (batch)Zhipu AI79
102GLM 5.2Zhipu AI79
103GLM 5.3 FlashZhipu AI78
104GLM 5.1Zhipu AI78
105GLM 5 TurboZhipu AI78
106GLM 5Zhipu AI78
107GLM 5.3 FlashXZhipu AI78
108GLM 5.3 Flash (batch)Zhipu AI78
109Qwen3.5-122B-A10BAlibaba78
110Qwen3.5-27BAlibaba77
111GPT-5 Mini (batch)OpenAI77
112GPT-5 Nano (batch)OpenAI77
113Gemini 3.5 Flash LiteGoogle77
114Gemini 3.5 Flash Lite (batch)Google77
115MiMo-V2.6-ProXiaomi76
116MiMo-V2.5-ProXiaomi76
117Qwen3.5-35B-A3BAlibaba76
118GLM 5.2 (free)Zhipu AI76
119Kimi K2.6Moonshot AI76
120Qwen3.8 Max (0902)Alibaba76
121Qwen3.7 PlusAlibaba76
122o3 MiniOpenAI75
123Claude Opus 4.1 (batch)Anthropic75
124GLM 4.7Zhipu AI75
125GLM 4.6Zhipu AI75
126GLM 4.5Zhipu AI75
127Qwen3.7 MaxAlibaba75
128Qwen3.5 Plus 2026-04-20Alibaba75
129Qwen3.6 Max PreviewAlibaba75
130Hy3Tencent74
131Claude Opus 4.1Anthropic74
132o1OpenAI74
133o3 Mini (batch)OpenAI74
134Qwen3.6 PlusAlibaba74
135MiniMax M3MiniMax74
136Claude Sonnet 4Anthropic74
137R1DeepSeek74
138o1-proOpenAI74
139Qwen3.8 27BAlibaba73
140MiMo-V2.5Xiaomi73
141Gemma 4 26B A4B Google73
142Gemma 4 26B A4B (free)Google73
143Qwen3.8 27B (free)Alibaba73
144Inklingthinkingmachines73
145Inkling (free)thinkingmachines73
146GLM 5V TurboZhipu AI72
147o4 Mini HighOpenAI72
148DeepSeek V3.2 ExpDeepSeek72
149DeepSeek V3.1DeepSeek72
150Mistral Medium 3.5Mistral AI72
151MiniMax M2.7MiniMax72
152MiniMax M2.5MiniMax72
153MiniMax M2.1MiniMax71
154MiniMax M2MiniMax71
155MiniMax M1MiniMax71
156GLM 4.5 AirZhipu AI71
157Claude Haiku 4.5Anthropic70
158Claude Haiku 4.5 (batch)Anthropic70
159Qwen3 VL 235B A22B ThinkingAlibaba69
160DeepSeek V3.1 TerminusDeepSeek69
161Inkling Smallthinkingmachines69
162Inkling Small (free)thinkingmachines69
163Qwen3.5-FlashAlibaba69
164Hy3 previewTencent68
165Qwen3.5 Plus 2026-02-15Alibaba68
166Qwen3 Max ThinkingAlibaba68
167Qwen3 Next 80B A3B ThinkingAlibaba67
168Qwen3.5-9BAlibaba67
169Step 3.5 FlashStepFun66
170Composer 2Cursor66
171Composer 2 FastCursor66
172GLM 4.6VZhipu AI66
173Qwen3 235B A22B Thinking 2507Alibaba65
174Qwen3 30B A3B Thinking 2507Alibaba64
175Qwen3 30B A3BAlibaba64
176o3 Mini HighOpenAI64
177GLM 4.7 FlashZhipu AI64
178Trinity Large Thinkingarcee-ai63
179GLM 4.5VZhipu AI62
180GPT-5 NanoOpenAI62
181gpt-oss-120bOpenAI62
182Mercury 2.5Inception61
183Qwen3 8BAlibaba61
184Mercury 2Inception61
185Nova 2 LiteAmazon60
186gpt-oss-20bOpenAI57
187gpt-oss-20b (batch)OpenAI57
188Qwen3 235B A22BAlibaba54
189Granite 4.2 8BIBM54
190GPT-5 MiniOpenAI54
191Kimi K2.5Moonshot AI48
192R1 Distill Llama 70BDeepSeek41
193GPT-6 Luna ProOpenAI40
194GPT-6 Luna Pro (batch)OpenAI40
195GPT-6 LunaOpenAI40
196GPT-6 Luna (batch)OpenAI40
197GPT-6 Sol ProOpenAI40
198GPT-6 Sol Pro (batch)OpenAI40
199GPT-6 SolOpenAI40
200GPT-6 Sol (batch)OpenAI40
201Claude Opus 5.5 (batch)Anthropic40
202MiMo-V2.6-Pro-UltraSpeedXiaomi40
203MiMo-V2.6-FlashXiaomi40
204Ternary Bonsai 2 27Bprism-ml40
205DeepSeek Pro Latest~deepseek40
206DeepSeek Flash Latest~deepseek40
207GPT Astra Latest~openai40
208GPT Sol Latest~openai40
209GPT Terra Latest~openai40
210GPT Luna Latest~openai40
211Fugu Ultra v2sakana40
212Fugu Maxsakana40
213Ling 3.0 Flash VLinclusionai40
214Ling 3.0 Flash VL (free)inclusionai40
215DeepSeek V4.1 FlashDeepSeek40
216DeepSeek V4.1 Flash (batch)DeepSeek40
217Ling 3.0 Flash Sante (free)inclusionai40
218Muse Spark 1.3 Contributormeta40
219Hy4 previewTencent40
220Ling 3.0 Flash Fininclusionai40
221Ling 3.0 Flash Fin (free)inclusionai40
222GLM Flash Latest~z-ai40
223Qwen3.8 FlashAlibaba40
224Muse Spark 1.2 Contributormeta40
225DeepSeek V4 Flash Vision ExpDeepSeek40
226GLM Latest~z-ai40
227Dots3-Note Preview (free)dots-studio40
228Gemini 3.7 FlashGoogle40
229Gemini 3.7 Flash (batch)Google40
230Seed 2.1 TurboByteDance40
231Qwen3.8 2.4T A95BAlibaba40
232Seed-2.0-CodeByteDance40
233DeepSeek V4 Pro 0813DeepSeek40
234LFM2.5-2.6B (free)Liquid AI40
235Nemotron 3.5 LightningNVIDIA40
236Nemotron 3.5 Lightning (free)NVIDIA40
237Sakana Namazusakana40
238Solar Pro 4Upstage40
239Muse Glimmer 30Bmeta40
240DeepSeek V4 Flash Latest~deepseek40
241DeepSeek V4 Flash 0731DeepSeek40
242Qwen3.7 FlashAlibaba40
243Claude Opus 5 (batch)Anthropic40
244Ling 3.0 Flashinclusionai40
245Laguna S 2.1poolside40
246Laguna S 2.1 (free)poolside40
247Gemini 3.6 FlashGoogle40
248Gemini 3.6 Flash (batch)Google40
249LongCat 2.0Meituan40
250Kimi K3Moonshot AI40
251Kimi K3 (batch)Moonshot AI40
252Grok Latest~x-ai40
253Aion-3.0-Miniaion-labs-
254Aion-3.0aion-labs-
255Laguna XS 2.1poolside-
256Laguna XS 2.1 (free)poolside-
257Claude Sonnet 5 (batch)Anthropic-
258Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google-
259Fugu Ultrasakana-
260Nano Banana 2 (Gemini 3.1 Flash Image)Google-
261Nano Banana Pro (Gemini 3 Pro Image)Google-
262North Mini Code (free)Cohere-
263Kimi K2.7 CodeMoonshot AI-
264Claude Fable Latest~anthropic-
265Nemotron 3.5 Content SafetyNVIDIA-
266Nemotron 3.5 Content Safety (free)NVIDIA-
267Nemotron 3 UltraNVIDIA-
268Nemotron 3 Ultra (free)NVIDIA-
269Step 3.7 FlashStepFun-
270Grok Build 0.1xAI-
271Perceptron Mk1perceptron-
272Gemini 3.1 Flash LiteGoogle-
273Gemini 3.1 Flash Lite (batch)Google-
274Mistral Medium 3.5 (batch)Mistral AI-
275Nemotron 3 Nano Omni (free)NVIDIA-
276Claude Haiku Latest~anthropic-
277GPT Mini Latest~openai-
278Gemini Pro Latest~google-
279Kimi Latest~moonshotai-
280Gemini Flash Latest~google-
281Claude Sonnet Latest~anthropic-
282Qwen3.6 FlashAlibaba-
283Qwen3.6 35B A3BAlibaba-
284Qwen3.6 27BAlibaba-
285DeepSeek V4 Pro 0423DeepSeek-
286DeepSeek V4 Flash 0423DeepSeek-
287GPT-5.4 Image 2OpenAI-
288Claude Opus Latest~anthropic-
289Mistral Small 4Mistral AI-
290Mistral Small 4 (batch)Mistral AI-
291Nemotron 3 SuperNVIDIA-
292Nemotron 3 Super (free)NVIDIA-
293Seed-2.0-LiteByteDance-
294Seed-2.0-MiniByteDance-
295Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google-
296Aion-2.0aion-labs-
297Solar Pro 3Upstage-
298Seed 1.6 FlashByteDance-
299Seed 1.6ByteDance-
300Nemotron 3 Nano 30B A3BNVIDIA-
301Nano Banana Pro (Gemini 3 Pro Image Preview)Google-
302Kimi K2 ThinkingMoonshot AI-
303Sonar Pro SearchPerplexity-
304gpt-oss-safeguard-20bOpenAI-
305GPT-5 Image MiniOpenAI-
306Qwen3 VL 8B ThinkingAlibaba-
307GPT-5 ImageOpenAI-
308Qwen3 VL 30B A3B ThinkingAlibaba-
309Hunyuan A13B InstructTencent-
310ERNIE 4.5 VL 424B A47B Baidu-
311Qwen3 14BAlibaba-
312Qwen3 32BAlibaba-
313Reka Flash 3rekaai-
314Sonar Reasoning ProPerplexity-
315Sonar Deep ResearchPerplexity-
316SWE-1.5Windsurf-
317Falcon-H1-Arabic 34B InstructTII-
318Falcon-H1-Arabic 7B InstructTII-
319Falcon-H1-Arabic 3B InstructTII-
320Falcon Arabic 7B InstructTII-
321Falcon3 10B InstructTII-
322Falcon3 7B InstructTII-

Top Standard Models (Non-Reasoning) - For Comparison

#ModelScore
1GPT-5.2 ChatOpenAI91
2Gemma 2 27BGoogle77
3GPT-4o (batch)OpenAI72
4DeepSeek V3 0324DeepSeek72
5GPT-4o (2024-11-20)OpenAI71
6GPT-4o (2024-08-06)OpenAI71
7GPT-4oOpenAI71
8GPT-4o (2024-05-13)OpenAI71
9Llama 4 MaverickMeta71
10DeepSeek V3DeepSeek70

Understanding AI Reasoning Benchmarks

What Is Chain-of-Thought?

Chain-of-thought (CoT) prompting enables AI models to break down complex problems into intermediate steps before producing a final answer. Models like OpenAI o1 and DeepSeek R1 internalize this process, generating hidden reasoning traces that dramatically improve accuracy on math, logic, and multi-step tasks compared to direct answering.

When Reasoning Helps

Reasoning models shine on tasks that require multiple logical steps: mathematical proofs, complex coding challenges, scientific analysis, strategic planning, and any problem where standard models tend to hallucinate or skip steps. For simple Q&A or creative writing, standard models are often faster and equally effective.

Speed vs Accuracy

Reasoning models consume more tokens and take longer to respond because they generate internal thinking traces. This trade-off is worthwhile when correctness matters more than latency - for example in code generation, financial analysis, or exam-style problems. For real-time chat, standard models remain the better choice.

Emerging Reasoning Models

The reasoning model landscape is evolving rapidly. OpenAI's o1 and o3 series led the way, followed by DeepSeek R1 bringing open-source reasoning. Google, Anthropic, and other providers have since introduced their own reasoning-capable models, driving down costs and expanding access to chain-of-thought capabilities.

Frequently Asked Questions

AI reasoning benchmarks test a model's ability to solve complex problems requiring logical thinking, mathematical reasoning, scientific analysis, and multi-step problem solving - tasks that go beyond simple pattern matching.

DeepSeek R1, OpenAI o3, and Claude with extended thinking lead on reasoning benchmarks. These models use chain-of-thought processing to break down complex problems into steps, achieving significantly higher accuracy.

Key reasoning benchmarks include GPQA Diamond (graduate-level science), MATH-500 (mathematical reasoning), AIME (competition math), ARC Challenge (science questions), and GSM8K (grade-school math). Each tests different aspects of reasoning ability.

AI Reasoning Benchmark - Which LLM Thinks Best? | LM Market Cap