Skip to content

Best AI Models for Coding

AI models ranked by coding ability using SWE-bench Verified, HumanEval, and BigCodeBench scores. Fallback to Arena Elo for unbenched models.

Last updated: 30m ago
#1 Model

Claude Opus 4.6

Score: 83.9

Average Score

58.7

Across all ranked models

Models Ranked

113

With benchmark data

Weights:SWE-bench Verified (40%)HumanEval (30%)BigCodeBench (30%)Fallback: Arena Elo

Top Best for Coding Models by Weighted Score

Top 15 models by weighted score

LMMarketCap.com

Benchmark Breakdown

Per-benchmark scores for top 10 models

SWE-bench Verified
HumanEval
BigCodeBench
LMMarketCap.com
#ModelScore
1Claude Opus 4.6Anthropic83.9
2Claude Sonnet 4.6Anthropic80.9
3GPT-5.4OpenAI78.7
4Claude Opus 4.5Anthropic78.3
5GPT-5.2OpenAI77.5
6GPT-5.1OpenAI76.7
7Claude Sonnet 4.5Anthropic76.2
8Claude Fable 5Anthropic76
9GPT-5OpenAI75.8
10Gemini 3 Flash PreviewGoogle75.6
11o3OpenAI74.3
12Claude Haiku 4.5Anthropic71.8
13o4 MiniOpenAI71.7
14GPT-5.5OpenAI71
15GPT-5.5 ProOpenAI71
16Claude Opus 4.8Anthropic70.9
17Claude Opus 4.7Anthropic70.1
18GPT-4o-miniOpenAI69.8
19Claude Sonnet 4Anthropic66.9
20Gemini 2.5 FlashGoogle65.8
21Gemini 3.1 Pro PreviewGoogle64.5
22Llama 4 MaverickMeta62.6
23GPT-4 TurboOpenAI60.9
24Llama 3.3 70B InstructMeta60.9
25GPT-4OpenAI60.5
26Muse Spark 1.1(fallback)meta60
27GPT-5.2 Chat(fallback)OpenAI60
28GLM 5.3 Flash(fallback)Zhipu AI60
29Grok 4.5(fallback)xAI60
30MiMo-V2.5-Pro(fallback)Xiaomi60
31GLM 5.1(fallback)Zhipu AI60
32Kimi K2.6(fallback)Moonshot AI60
33Qwen3.6 Max Preview(fallback)Alibaba60
34GLM 5(fallback)Zhipu AI60
35Gemini 3.5 Flash Lite(fallback)Google60
36Hy3(fallback)Tencent60
37Qwen3.7 Plus(fallback)Alibaba60
38Gemma 4 31B(fallback)Google60
39Claude Opus 4.1(fallback)Anthropic60
40Qwen3.6 Plus(fallback)Alibaba60
41Qwen3.5 397B A17B(fallback)Alibaba60
42GLM 4.7(fallback)Zhipu AI60
43MiniMax M3(fallback)MiniMax60
44Inkling(fallback)thinkingmachines60
45Gemma 4 26B A4B (fallback)Google60
46Qwen3.8 27B(fallback)Alibaba60
47MiMo-V2.5(fallback)Xiaomi60
48GLM 5V Turbo(fallback)Zhipu AI60
49Gemini 3.1 Flash Lite Preview(fallback)Google60
50Mistral Medium 3.5(fallback)Mistral AI60
51DeepSeek V3.2(fallback)DeepSeek60
52GLM 4.6(fallback)Zhipu AI60
53DeepSeek V3.2 Exp(fallback)DeepSeek60
54DeepSeek V3.1(fallback)DeepSeek60
55Qwen3.5-122B-A10B(fallback)Alibaba60
56MiniMax M2.7(fallback)MiniMax60
57DeepSeek V3.1 Terminus(fallback)DeepSeek60
58Qwen3 VL 235B A22B Instruct(fallback)Alibaba60
59Hy3 preview(fallback)Tencent60
60GLM 4.5(fallback)Zhipu AI60
61Qwen3.5-27B(fallback)Alibaba60
62Inkling Small(fallback)thinkingmachines60
63Qwen3 Next 80B A3B Instruct(fallback)Alibaba60
64Qwen3.5-Flash(fallback)Alibaba60
65Qwen3 VL 235B A22B Thinking(fallback)Alibaba60
66Qwen3.5-35B-A3B(fallback)Alibaba60
67Step 3.5 Flash(fallback)StepFun60
68MiniMax M2.5(fallback)MiniMax60
69GPT-5 Mini(fallback)OpenAI60
70GLM 4.6V(fallback)Zhipu AI60
71GLM 4.5 Air(fallback)Zhipu AI60
72Qwen3 Next 80B A3B Thinking(fallback)Alibaba60
73Trinity Large Thinking(fallback)arcee-ai60
74GLM 4.7 Flash(fallback)Zhipu AI60
75MiniMax M1(fallback)MiniMax60
76o3 Mini High(fallback)OpenAI60
77Command A(fallback)Cohere60
78GLM 4.5V(fallback)Zhipu AI60
79gpt-oss-120b(fallback)OpenAI60
80Qwen3 8B(fallback)Alibaba60
81Mercury 2(fallback)Inception60
82MiniMax M2(fallback)MiniMax60
83GPT-5 Nano(fallback)OpenAI60
84Nova 2 Lite(fallback)Amazon60
85gpt-oss-20b(fallback)OpenAI60
86Mistral Large 2407(fallback)Mistral AI60
87Granite 4.2 8B(fallback)IBM60
88Gemini 3.5 Flash(fallback)Google60
89GPT-4.1OpenAI58.8
90Phi 4Microsoft57.6
91Llama 3.1 70B InstructMeta57
92o1OpenAI57
93DeepSeek V3DeepSeek56.6
94Gemma 2 27BGoogle55.6
95Mistral LargeMistral AI54.9
96GPT-4oOpenAI54.7
97Claude 3 HaikuAnthropic52.3
98DeepSeek V3 0324DeepSeek50.5
99R1 0528DeepSeek46.1
100Llama 3.1 8B InstructMeta46
101Gemini 2.5 ProGoogle44.3
102Llama 4 ScoutMeta40.9
103GPT-4.1 MiniOpenAI39.1
104GPT-4o (2024-11-20)OpenAI38.4
105o3 MiniOpenAI38.1
106GPT-4o-mini (2024-07-18)OpenAI36.9
107R1DeepSeek36.8
108Qwen2.5 7B InstructAlibaba30.1
109Command R+ (08-2024)Cohere29.7
110R1 Distill Llama 70BDeepSeek28.2
111GPT-4.1 NanoOpenAI22.7
112Llama 3.2 3B InstructMeta18.7
113Llama 3.2 1B InstructMeta6.6

How scores are calculated

Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. Models without any primary benchmark data fall back to Arena Elo (normalized to 0-100) and are marked accordingly. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.

Frequently Asked Questions

Based on our benchmark analysis, Claude Opus 4.6 by Anthropic is currently the #1 ranked model for coding, with a weighted score of 83.9/100.

Models are ranked using a weighted average of SWE-bench Verified, HumanEval, BigCodeBench benchmark scores. Models without primary benchmark data fall back to Arena Elo. All scores are normalized to a 0-100 scale.

We currently rank 113 models that have relevant benchmark data for coding tasks.

Best AI Models for Coding (2026) | LM Market Cap