Skip to content

AI Reasoning Benchmark

How do reasoning models stack up against standard LLMs? This benchmark compares 324 reasoning models against 116 standard models on composite score, pricing, and capabilities - helping you decide when chain-of-thought thinking is worth the trade-off.

Reasoning vs Standard - Head-to-Head

Reasoning Models
324
Models
96
Top Score
70
Avg Score
$10.77
Avg $/1M Out
Standard Models
116
Models
91
Top Score
61
Avg Score
$3.23
Avg $/1M Out

Reasoning models from 45 providers. Score difference: +9 points average for reasoning models.

Reasoning Models - Ranked by Score

#ModelScore
1Claude Fable 5.1Anthropic96
2Claude Fable 5.1 (batch)Anthropic96
3Claude Fable 5Anthropic96
4Claude Fable 5 (batch)Anthropic96
5Claude Opus 5.5Anthropic95
6Claude Opus 5Anthropic95
7Claude Opus 4.8Anthropic95
8Claude Opus 4.7Anthropic95
9Claude Opus 4.7 (batch)Anthropic95
10Claude Opus 4.8 (batch)Anthropic95
11GPT-5.5 ProOpenAI93
12GPT-5.5 Pro (batch)OpenAI93
13GPT-5.5OpenAI93
14GPT-5.5 (batch)OpenAI93
15Gemini 3.1 Pro Preview Custom ToolsGoogle92
16Gemini 3.1 Pro PreviewGoogle92
17Gemini 3.1 Pro Preview (batch)Google92
18GPT-5.4 ProOpenAI92
19GPT-5.4 Pro (batch)OpenAI92
20GPT-5.4OpenAI92
21GPT-5.4 (batch)OpenAI92
22GPT-5.3-CodexOpenAI91
23GPT-5.2-CodexOpenAI91
24GPT-5.2 ProOpenAI91
25GPT-5.2 Pro (batch)OpenAI91
26GPT-5.2OpenAI91
27GPT-5.2 (batch)OpenAI91
28Claude Opus 4.6Anthropic90
29Claude Opus 4.6 (batch)Anthropic90
30GPT-5.6 Luna ProOpenAI89
31GPT-5.6 Luna Pro (batch)OpenAI89
32GPT-5.6 LunaOpenAI89
33GPT-5.6 Luna (batch)OpenAI89
34GPT-5.6 Terra ProOpenAI89
35GPT-5.6 Terra Pro (batch)OpenAI89
36GPT-5.6 TerraOpenAI89
37GPT-5.6 Terra (batch)OpenAI89
38GPT-5.6 Sol ProOpenAI89
39GPT-5.6 Sol Pro (batch)OpenAI89
40GPT-5.6 SolOpenAI89
41GPT-5.6 Sol (batch)OpenAI89
42Grok 4.7xAI89
43Grok 4.6xAI89
44Grok 4.5xAI89
45Grok 4.3xAI89
46Grok 4.20xAI89
47GPT-5.1-Codex-MaxOpenAI89
48GPT-5.1OpenAI89
49GPT-5.1-CodexOpenAI89
50GPT-5 ProOpenAI89
51GPT-5 Pro (batch)OpenAI89
52GPT-5OpenAI89
53GPT-5 (batch)OpenAI89
54Gemini 3 Flash PreviewGoogle88
55Gemini 3 Flash Preview (batch)Google88
56Grok 4.20 Multi-AgentxAI88
57GPT-5.1 (batch)OpenAI88
58GPT-5.1-Codex-MiniOpenAI88
59o3 ProOpenAI87
60o3OpenAI87
61o3 (batch)OpenAI87
62Claude Sonnet 5Anthropic85
63Claude Sonnet 4.6Anthropic85
64Claude Sonnet 4.6 (batch)Anthropic85
65Claude Opus 4.5Anthropic85
66Claude Opus 4.5 (batch)Anthropic85
67Gemini 2.5 ProGoogle84
68Gemini 2.5 Pro (batch)Google84
69Gemini 2.5 Pro Preview 06-05Google84
70DeepSeek V3.2DeepSeek83
71Claude Sonnet 4.5Anthropic82
72Claude Sonnet 4.5 (batch)Anthropic82
73GPT-6 AstraOpenAI82
74GPT-6 Astra (batch)OpenAI82
75GPT-6 Astra ProOpenAI82
76GPT-6 Astra Pro (batch)OpenAI82
77o4 MiniOpenAI81
78o4 Mini (batch)OpenAI81
79Gemini 3.8 FlashGoogle81
80Gemini 3.8 Flash (batch)Google81
81Grok 4.3 (batch)xAI81
82Muse Spark 1.3meta81
83Muse Spark 1.2meta81
84Muse Spark 1.1meta81
85Gemma 4 31BGoogle81
86Gemma 4 31B (free)Google81
87Qwen3.5 397B A17BAlibaba79
88R1 0528DeepSeek79
89GPT-5.4 NanoOpenAI79
90GPT-5.4 Nano (batch)OpenAI79
91GPT-5.4 MiniOpenAI79
92GPT-5.4 Mini (batch)OpenAI79
93Gemini 3.1 Flash Lite PreviewGoogle79
94Gemini 2.5 Flash LiteGoogle79
95Gemini 2.5 Flash Lite (batch)Google79
96Gemini 2.5 FlashGoogle79
97Gemini 2.5 Flash (batch)Google79
98Gemini 3.5 FlashGoogle79
99Gemini 3.5 Flash (batch)Google79
100GLM 5.3Zhipu AI79
101GLM 5.3 (batch)Zhipu AI79
102GLM 5.2Zhipu AI79
103GLM 5.3 FlashZhipu AI78
104GLM 5.1Zhipu AI78
105GLM 5 TurboZhipu AI78
106GLM 5Zhipu AI78
107GLM 5.3 FlashXZhipu AI78
108GLM 5.3 Flash (batch)Zhipu AI78
109Qwen3.5-122B-A10BAlibaba78
110Qwen3.5-27BAlibaba77
111GPT-5 Mini (batch)OpenAI77
112GPT-5 Nano (batch)OpenAI77
113Gemini 3.5 Flash LiteGoogle77
114Gemini 3.5 Flash Lite (batch)Google77
115MiMo-V2.6-ProXiaomi76
116MiMo-V2.5-ProXiaomi76
117Qwen3.5-35B-A3BAlibaba76
118GLM 5.2 (free)Zhipu AI76
119Kimi K2.6Moonshot AI76
120Qwen3.8 Max (0902)Alibaba76
121Qwen3.7 PlusAlibaba76
122o3 MiniOpenAI75
123Claude Opus 4.1 (batch)Anthropic75
124GLM 4.7Zhipu AI75
125GLM 4.6Zhipu AI75
126GLM 4.5Zhipu AI75
127Qwen3.7 MaxAlibaba75
128Qwen3.5 Plus 2026-04-20Alibaba75
129Qwen3.6 Max PreviewAlibaba75
130Hy3Tencent74
131Claude Opus 4.1Anthropic74
132o1OpenAI74
133o3 Mini (batch)OpenAI74
134Qwen3.6 PlusAlibaba74
135MiniMax M3MiniMax74
136Claude Sonnet 4Anthropic74
137R1DeepSeek74
138o1-proOpenAI74
139Qwen3.8 27BAlibaba73
140MiMo-V2.5Xiaomi73
141Gemma 4 26B A4B Google73
142Gemma 4 26B A4B (free)Google73
143Qwen3.8 27B (free)Alibaba73
144Inklingthinkingmachines73
145Inkling (free)thinkingmachines73
146GLM 5V TurboZhipu AI72
147o4 Mini HighOpenAI72
148DeepSeek V3.2 ExpDeepSeek72
149DeepSeek V3.1DeepSeek72
150Mistral Medium 3.5Mistral AI72
151MiniMax M2.7MiniMax72
152MiniMax M2.5MiniMax72
153MiniMax M2.1MiniMax71
154MiniMax M2MiniMax71
155MiniMax M1MiniMax71
156GLM 4.5 AirZhipu AI71
157Claude Haiku 4.5Anthropic70
158Claude Haiku 4.5 (batch)Anthropic70
159Qwen3 VL 235B A22B ThinkingAlibaba69
160DeepSeek V3.1 TerminusDeepSeek69
161Inkling Smallthinkingmachines69
162Inkling Small (free)thinkingmachines69
163Qwen3.5-FlashAlibaba69
164Hy3 previewTencent68
165Qwen3.5 Plus 2026-02-15Alibaba68
166Qwen3 Max ThinkingAlibaba68
167Qwen3 Next 80B A3B ThinkingAlibaba67
168Qwen3.5-9BAlibaba67
169Step 3.5 FlashStepFun66
170Composer 2Cursor66
171Composer 2 FastCursor66
172GLM 4.6VZhipu AI66
173Qwen3 235B A22B Thinking 2507Alibaba65
174Qwen3 30B A3B Thinking 2507Alibaba64
175Qwen3 30B A3BAlibaba64
176o3 Mini HighOpenAI64
177GLM 4.7 FlashZhipu AI64
178Trinity Large Thinkingarcee-ai63
179GLM 4.5VZhipu AI62
180GPT-5 NanoOpenAI62
181gpt-oss-120bOpenAI62
182Mercury 2.5Inception61
183Qwen3 8BAlibaba61
184Mercury 2Inception61
185Nova 2 LiteAmazon60
186gpt-oss-20bOpenAI57
187gpt-oss-20b (batch)OpenAI57
188Qwen3 235B A22BAlibaba54
189Granite 4.2 8BIBM54
190GPT-5 MiniOpenAI54
191Command A+Cohere51
192Kimi K2.5Moonshot AI48
193R1 Distill Llama 70BDeepSeek41
194GPT-6 Luna ProOpenAI40
195GPT-6 Luna Pro (batch)OpenAI40
196GPT-6 LunaOpenAI40
197GPT-6 Luna (batch)OpenAI40
198GPT-6 Sol ProOpenAI40
199GPT-6 Sol Pro (batch)OpenAI40
200GPT-6 SolOpenAI40
201GPT-6 Sol (batch)OpenAI40
202Claude Opus 5.5 (batch)Anthropic40
203MiMo-V2.6-Pro-UltraSpeedXiaomi40
204MiMo-V2.6-FlashXiaomi40
205Qwen3.8 Omni FlashAlibaba40
206Ternary Bonsai 2 27Bprism-ml40
207DeepSeek Pro Latest~deepseek40
208DeepSeek Flash Latest~deepseek40
209GPT Astra Latest~openai40
210GPT Sol Latest~openai40
211GPT Terra Latest~openai40
212GPT Luna Latest~openai40
213Fugu Ultra v2sakana40
214Fugu Maxsakana40
215Ling 3.0 Flash VLinclusionai40
216Ling 3.0 Flash VL (free)inclusionai40
217DeepSeek V4.1 FlashDeepSeek40
218DeepSeek V4.1 Flash (batch)DeepSeek40
219Ling 3.0 Flash Sante (free)inclusionai40
220Muse Spark 1.3 Contributormeta40
221Hy4 previewTencent40
222Ling 3.0 Flash Fininclusionai40
223Ling 3.0 Flash Fin (free)inclusionai40
224GLM Flash Latest~z-ai40
225Qwen3.8 FlashAlibaba40
226Muse Spark 1.2 Contributormeta40
227DeepSeek V4 Flash Vision ExpDeepSeek40
228GLM Latest~z-ai40
229Dots3-Note Preview (free)dots-studio40
230Gemini 3.7 FlashGoogle40
231Gemini 3.7 Flash (batch)Google40
232Seed 2.1 TurboByteDance40
233Qwen3.8 2.4T A95BAlibaba40
234Seed-2.0-CodeByteDance40
235DeepSeek V4 Pro 0813DeepSeek40
236LFM2.5-2.6B (free)Liquid AI40
237Nemotron 3.5 LightningNVIDIA40
238Nemotron 3.5 Lightning (free)NVIDIA40
239Sakana Namazusakana40
240Solar Pro 4Upstage40
241Muse Glimmer 30Bmeta40
242DeepSeek V4 Flash Latest~deepseek40
243DeepSeek V4 Flash 0731DeepSeek40
244Qwen3.7 FlashAlibaba40
245Claude Opus 5 (batch)Anthropic40
246Ling 3.0 Flashinclusionai40
247Laguna S 2.1poolside40
248Laguna S 2.1 (free)poolside40
249Gemini 3.6 FlashGoogle40
250Gemini 3.6 Flash (batch)Google40
251LongCat 2.0Meituan40
252Kimi K3Moonshot AI40
253Kimi K3 (batch)Moonshot AI40
254Grok Latest~x-ai-
255Aion-3.0-Miniaion-labs-
256Aion-3.0aion-labs-
257Laguna XS 2.1poolside-
258Laguna XS 2.1 (free)poolside-
259Claude Sonnet 5 (batch)Anthropic-
260Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google-
261Fugu Ultrasakana-
262Nano Banana 2 (Gemini 3.1 Flash Image)Google-
263Nano Banana Pro (Gemini 3 Pro Image)Google-
264North Mini Code (free)Cohere-
265Kimi K2.7 CodeMoonshot AI-
266Claude Fable Latest~anthropic-
267Nemotron 3.5 Content SafetyNVIDIA-
268Nemotron 3.5 Content Safety (free)NVIDIA-
269Nemotron 3 UltraNVIDIA-
270Nemotron 3 Ultra (free)NVIDIA-
271Step 3.7 FlashStepFun-
272Grok Build 0.1xAI-
273Perceptron Mk1perceptron-
274Gemini 3.1 Flash LiteGoogle-
275Gemini 3.1 Flash Lite (batch)Google-
276Mistral Medium 3.5 (batch)Mistral AI-
277Nemotron 3 Nano Omni (free)NVIDIA-
278Claude Haiku Latest~anthropic-
279GPT Mini Latest~openai-
280Gemini Pro Latest~google-
281Kimi Latest~moonshotai-
282Gemini Flash Latest~google-
283Claude Sonnet Latest~anthropic-
284Qwen3.6 FlashAlibaba-
285Qwen3.6 35B A3BAlibaba-
286Qwen3.6 27BAlibaba-
287DeepSeek V4 Pro 0423DeepSeek-
288DeepSeek V4 Flash 0423DeepSeek-
289GPT-5.4 Image 2OpenAI-
290Claude Opus Latest~anthropic-
291Mistral Small 4Mistral AI-
292Mistral Small 4 (batch)Mistral AI-
293Nemotron 3 SuperNVIDIA-
294Nemotron 3 Super (free)NVIDIA-
295Seed-2.0-LiteByteDance-
296Seed-2.0-MiniByteDance-
297Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google-
298Aion-2.0aion-labs-
299Solar Pro 3Upstage-
300Seed 1.6 FlashByteDance-
301Seed 1.6ByteDance-
302Nemotron 3 Nano 30B A3BNVIDIA-
303Nano Banana Pro (Gemini 3 Pro Image Preview)Google-
304Kimi K2 ThinkingMoonshot AI-
305Sonar Pro SearchPerplexity-
306gpt-oss-safeguard-20bOpenAI-
307GPT-5 Image MiniOpenAI-
308Qwen3 VL 8B ThinkingAlibaba-
309GPT-5 ImageOpenAI-
310Qwen3 VL 30B A3B ThinkingAlibaba-
311Hunyuan A13B InstructTencent-
312ERNIE 4.5 VL 424B A47B Baidu-
313Qwen3 14BAlibaba-
314Qwen3 32BAlibaba-
315Reka Flash 3rekaai-
316Sonar Reasoning ProPerplexity-
317Sonar Deep ResearchPerplexity-
318SWE-1.5Windsurf-
319Falcon-H1-Arabic 34B InstructTII-
320Falcon-H1-Arabic 7B InstructTII-
321Falcon-H1-Arabic 3B InstructTII-
322Falcon Arabic 7B InstructTII-
323Falcon3 10B InstructTII-
324Falcon3 7B InstructTII-

Top Standard Models (Non-Reasoning) - For Comparison

#ModelScore
1GPT-5.2 ChatOpenAI91
2Gemma 2 27BGoogle77
3GPT-4o (batch)OpenAI72
4DeepSeek V3 0324DeepSeek72
5GPT-4o (2024-11-20)OpenAI71
6GPT-4o (2024-08-06)OpenAI71
7GPT-4oOpenAI71
8GPT-4o (2024-05-13)OpenAI71
9Llama 4 MaverickMeta71
10DeepSeek V3DeepSeek70

Understanding AI Reasoning Benchmarks

What Is Chain-of-Thought?

Chain-of-thought (CoT) prompting enables AI models to break down complex problems into intermediate steps before producing a final answer. Models like OpenAI o1 and DeepSeek R1 internalize this process, generating hidden reasoning traces that dramatically improve accuracy on math, logic, and multi-step tasks compared to direct answering.

When Reasoning Helps

Reasoning models shine on tasks that require multiple logical steps: mathematical proofs, complex coding challenges, scientific analysis, strategic planning, and any problem where standard models tend to hallucinate or skip steps. For simple Q&A or creative writing, standard models are often faster and equally effective.

Speed vs Accuracy

Reasoning models consume more tokens and take longer to respond because they generate internal thinking traces. This trade-off is worthwhile when correctness matters more than latency - for example in code generation, financial analysis, or exam-style problems. For real-time chat, standard models remain the better choice.

Emerging Reasoning Models

The reasoning model landscape is evolving rapidly. OpenAI's o1 and o3 series led the way, followed by DeepSeek R1 bringing open-source reasoning. Google, Anthropic, and other providers have since introduced their own reasoning-capable models, driving down costs and expanding access to chain-of-thought capabilities.

Frequently Asked Questions

AI reasoning benchmarks test a model's ability to solve complex problems requiring logical thinking, mathematical reasoning, scientific analysis, and multi-step problem solving - tasks that go beyond simple pattern matching.

DeepSeek R1, OpenAI o3, and Claude with extended thinking lead on reasoning benchmarks. These models use chain-of-thought processing to break down complex problems into steps, achieving significantly higher accuracy.

Key reasoning benchmarks include GPQA Diamond (graduate-level science), MATH-500 (mathematical reasoning), AIME (competition math), ARC Challenge (science questions), and GSM8K (grade-school math). Each tests different aspects of reasoning ability.

AI Reasoning Benchmark - Which LLM Thinks Best? | LM Market Cap