Skip to content

Best AI Models for Coding

AI models ranked by coding ability using SWE-bench Verified, HumanEval, and BigCodeBench scores. Fallback to Arena Elo for unbenched models.

Last updated: 36m ago
#1 Model

Claude Opus 4.6

Score: 83.9

Average Score

57.4

Across all ranked models

Models Ranked

118

With benchmark data

Weights:SWE-bench Verified (40%)HumanEval (30%)BigCodeBench (30%)Fallback: Arena Elo

Top Best for Coding Models by Weighted Score

Top 15 models by weighted score

LMMarketCap.com

Benchmark Breakdown

Per-benchmark scores for top 10 models

SWE-bench Verified
HumanEval
BigCodeBench
LMMarketCap.com
#ModelScore
1Claude Opus 4.6Anthropic83.9
2Claude Sonnet 4.6Anthropic80.9
3GPT-5.4OpenAI78.7
4Claude Opus 4.5Anthropic78.3
5GPT-5.2OpenAI77.5
6GPT-5.1OpenAI76.7
7Claude Sonnet 4.5Anthropic76.2
8Claude Fable 5Anthropic76
9GPT-5OpenAI75.8
10Gemini 3 Flash PreviewGoogle75.6
11o3OpenAI74.3
12Claude Opus 4Anthropic73.9
13o4 MiniOpenAI71.7
14GPT-5.5OpenAI71
15GPT-5.5 ProOpenAI71
16Claude Opus 4.8Anthropic70.9
17Claude Opus 4.7Anthropic70.1
18GPT-4o-miniOpenAI69.8
19Claude Haiku 4.5Anthropic68.9
20Claude Sonnet 4Anthropic66.9
21Gemini 2.5 FlashGoogle65.8
22Gemini 3.1 Pro PreviewGoogle64.5
23DeepSeek V4 ProDeepSeek64.5
24GPT-4 TurboOpenAI60.9
25Llama 3.3 70B InstructMeta60.9
26MiniMax M2.5MiniMax60.6
27GPT-4OpenAI60.5
28Qwen3.8 Max(fallback)Alibaba60
29Muse Spark 1.1(fallback)meta60
30Gemini 3.6 Flash(fallback)Google60
31GPT-5.2 Chat(fallback)OpenAI60
32Grok 4.5(fallback)xAI60
33GLM 5.1(fallback)Zhipu AI60
34MiMo-V2.5-Pro(fallback)Xiaomi60
35Kimi K2.6(fallback)Moonshot AI60
36Qwen3.6 Max Preview(fallback)Alibaba60
37Gemini 3.5 Flash Lite(fallback)Google60
38Qwen3.7 Plus(fallback)Alibaba60
39Hy3(fallback)Tencent60
40Gemma 4 31B(fallback)Google60
41Claude Opus 4.1(fallback)Anthropic60
42MiniMax M3(fallback)MiniMax60
43Qwen3.6 Plus(fallback)Alibaba60
44Qwen3.5 397B A17B(fallback)Alibaba60
45GLM 4.7(fallback)Zhipu AI60
46Inkling(fallback)thinkingmachines60
47Gemma 4 26B A4B (fallback)Google60
48MiMo-V2.5(fallback)Xiaomi60
49GLM 5V Turbo(fallback)Zhipu AI60
50Gemini 3.1 Flash Lite Preview(fallback)Google60
51Mistral Medium 3.5(fallback)Mistral AI60
52DeepSeek V3.2 Exp(fallback)DeepSeek60
53DeepSeek V3.1(fallback)DeepSeek60
54Qwen3.5-122B-A10B(fallback)Alibaba60
55MiniMax M2.7(fallback)MiniMax60
56Qwen3 VL 235B A22B Instruct(fallback)Alibaba60
57DeepSeek V3.1 Terminus(fallback)DeepSeek60
58Hy3 preview(fallback)Tencent60
59Qwen3.5-27B(fallback)Alibaba60
60Qwen3 Next 80B A3B Instruct(fallback)Alibaba60
61Qwen3.5-Flash(fallback)Alibaba60
62Qwen3 VL 235B A22B Thinking(fallback)Alibaba60
63Qwen3.5-35B-A3B(fallback)Alibaba60
64Step 3.5 Flash(fallback)StepFun60
65GLM 4.6V(fallback)Zhipu AI60
66GLM 4.5 Air(fallback)Zhipu AI60
67Qwen3 Next 80B A3B Thinking(fallback)Alibaba60
68Trinity Large Thinking(fallback)arcee-ai60
69GLM 4.7 Flash(fallback)Zhipu AI60
70MiniMax M1(fallback)MiniMax60
71o3 Mini High(fallback)OpenAI60
72Command A(fallback)Cohere60
73GLM 4.5V(fallback)Zhipu AI60
74Qwen3 8B(fallback)Alibaba60
75Mercury 2(fallback)Inception60
76Nova 2 Lite(fallback)Amazon60
77gpt-oss-20b(fallback)OpenAI60
78Mistral Large 2407(fallback)Mistral AI60
79Granite 4.1 8B(fallback)IBM60
80Olmo 3 32B Think(fallback)Allen AI60
81Inkling Small(fallback)thinkingmachines60
82Gemini 3.5 Flash(fallback)Google60
83GPT-4.1OpenAI58.8
84GLM 5Zhipu AI58.2
85GPT-5.2-CodexOpenAI58.2
86Phi 4Microsoft57.6
87Llama 3.1 70B InstructMeta57
88o1OpenAI57
89DeepSeek V3DeepSeek56.6
90DeepSeek V3.2DeepSeek56
91Gemma 2 27BGoogle55.6
92Mistral LargeMistral AI54.9
93GPT-4oOpenAI54.7
94GPT-5.1-CodexOpenAI52.8
95Claude 3 HaikuAnthropic52.3
96DeepSeek V3 0324DeepSeek50.5
97Llama 4 MaverickMeta50.2
98MiniMax M2MiniMax48.8
99GPT-5 MiniOpenAI47.8
100R1 0528DeepSeek46.1
101Llama 3.1 8B InstructMeta46
102Gemini 2.5 ProGoogle44.3
103GLM 4.6Zhipu AI44.3
104GLM 4.5Zhipu AI43.4
105GPT-4o (2024-11-20)OpenAI38.4
106o3 MiniOpenAI38.1
107GPT-4o-mini (2024-07-18)OpenAI36.9
108R1DeepSeek36.8
109GPT-4.1 MiniOpenAI31.2
110Llama 4 ScoutMeta30.9
111Qwen2.5 7B InstructAlibaba30.1
112Command R+ (08-2024)Cohere29.7
113R1 Distill Llama 70BDeepSeek28.2
114GPT-5 NanoOpenAI27.8
115GPT-4.1 NanoOpenAI22.7
116gpt-oss-120bOpenAI20.8
117Llama 3.2 3B InstructMeta18.7
118Llama 3.2 1B InstructMeta6.6

How scores are calculated

Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. Models without any primary benchmark data fall back to Arena Elo (normalized to 0-100) and are marked accordingly. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.

Frequently Asked Questions

Based on our benchmark analysis, Claude Opus 4.6 by Anthropic is currently the #1 ranked model for coding, with a weighted score of 83.9/100.

Models are ranked using a weighted average of SWE-bench Verified, HumanEval, BigCodeBench benchmark scores. Models without primary benchmark data fall back to Arena Elo. All scores are normalized to a 0-100 scale.

We currently rank 118 models that have relevant benchmark data for coding tasks.

Best AI Models for Coding (2026) | LM Market Cap