Skip to content
基准测试类别

Coding 基准测试

Compare top models across the benchmark suite that best represents coding performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.

6

类别中的基准测试

85

有覆盖的模型

0

有人类基准的基准测试

1

饱和的基准测试

包含内容

The current benchmark set in this category, with context on what each test captures.

所有基准测试

HumanEval Code Generation

饱和

The most recognized coding benchmark, though becoming saturated above 90%. Evidence of training data contamination in some models.

pass@1 %

Software Engineering Benchmark (Verified)

The gold standard for real-world coding ability. Unlike HumanEval, tests understanding of large codebases, debugging, and complex changes. Scores range 20-80%.

resolve rate %

BigCodeBench (Hard)

More realistic than HumanEval — tests practical programming skills including library usage, API calls, and multi-file reasoning.

pass@1 %

CursorBench (Multi-file Editing)

The first benchmark designed specifically for agentic coding assistants that edit multiple files. More realistic than single-function benchmarks like HumanEval.

pass rate %

Terminal-Bench (Terminal Agent Tasks)

Measures agentic capability in terminal environments — critical for AI coding assistants that execute commands and manage development workflows.

pass rate %

SWE-bench Multilingual

Most real codebases are polyglot. This benchmark tests whether coding models can handle the diversity of languages seen in production software engineering.

resolved rate %
Model type:

HumanEval Code Generation

164 Python function-generation problems where models must write correct code from docstrings, tested against unit tests. The original code generation benchmark.

Metric: pass@1 %Saturated

Why it matters

The most recognized coding benchmark, though becoming saturated above 90%. Evidence of training data contamination in some models.

HumanEval Scores (50 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇GPT-5.497.5%
2🥈o397.0%
3🥉GPT-5.297.0%
4GPT-5.196.8%
5GPT-596.5%
6o1 Preview96.3%
7Claude Opus 4.696.0%
8Gemini 3 Pro96.0%
9Grok 495.5%
10Claude Opus 4.595.2%
11Claude Sonnet 4.695.2%
12Claude Opus 495.0%
13o4-mini95.0%
14Claude Sonnet 4.594.5%
15Claude 3.7 Sonnet94.0%
16Claude Sonnet 493.8%
17Claude 3.5 Sonnet93.7%
18Qwen 2.5 Coder 32B92.7%
19o192.4%
20o1-mini92.4%
21Mistral Large 292.0%
22Mistral Large 292.0%
23Gemini 3 Flash92.0%
24GPT-4.191.5%
25Grok 390.5%
26GPT-4o90.2%
27Gemini 2.5 Flash90.0%
28Claude Haiku 4.589.8%
29Llama 4 Maverick89.5%
30Gemini 2.0 Flash89.4%
31Llama 3.1 405B89.0%
32Llama 3.3 70B88.4%
33GPT-488.4%
34Claude 3.5 Haiku88.1%
35GPT-4o mini87.2%
36GPT-4 Turbo87.1%
37Qwen 2.5 72B86.6%
38Claude 3 Opus84.9%
39DeepSeek V3 (March 2025)84.5%
40Gemini 1.5 Pro84.1%
41Grok 284.0%
42DeepSeek V382.6%
43Phi-482.6%
44Llama 3.1 70B80.5%
45Mixtral 8x22B78.6%
46Claude 3 Haiku76.8%
47Command R+74.3%
48Llama 4 Scout74.1%
49Gemma 2 27B69.5%
50Llama 3.1 8B Instruct69.5%

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-08T06:30:09.792Z. Some scores are approximate.

Frequently Asked Questions

AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.

Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.

A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.

Coding AI Benchmarks - Compare Top Models | LM Market Cap