Skip to content
最近更新: 2h ago
Coding 基准测试

Software Engineering Benchmark (Verified) 排行榜

Can a model resolve real GitHub issues from popular Python repositories? Human-validated subset ensures accurate evaluation. Tests end-to-end software engineering ability.

为什么重要: The gold standard for real-world coding ability. Unlike HumanEval, tests understanding of large codebases, debugging, and complex changes. Scores range 20-80%.

顶级模型

95%

Claude Fable 5

平均评分

62.0%

共52个模型

已测试模型

52

指标: resolve rate

人类基准

-

评分范围: 0%100%

SWE-bench Verified Scores - Top 25 Models

Ranked by SWE-bench Verified score (%)

LMMarketCap.com

模型排名

All models with a reported SWE-bench Verified score, ranked by highest resolve rate.

#2
88.7%
#2
88.7%
#8
80.6%
#10
80%
#12
78%
#15
76.5%
#16
76.2%
#17
75.8%
#18
75%
#19
72.8%
#20
72.7%
#21
72.5%
#22
71.7%
#23
70.8%
#25
70%
#27
68.1%
#29
63.8%
#30
63.4%
#31
61%
#32
60.4%
#33
59.8%
#34
57.6%
#35
56.4%
#36
55.4%
#36
55.4%
#38
54.6%
#39
53.8%
#40
49.3%
#41
49.2%
#43
48.9%
#45
34.8%
#46
30.8%
#47
26%
#48
23.9%
#50
13.5%
#51
9.1%

关于 SWE-bench Verified

全名
Software Engineering Benchmark (Verified)
类别
Coding
指标
resolve rate (%)
评分范围
0%100%
人类基准
尚未确定
状态
启用
Frequently Asked Questions

SWE-bench Verified is a standardized evaluation that measures AI model performance on specific tasks. It provides comparable scores across different models, helping developers choose the right model for their needs.

Claude Fable 5 currently holds the top score on the SWE-bench Verified benchmark. See our full rankings table above for the complete leaderboard with 52 models.

We update benchmark data from multiple sources including HuggingFace open-source model leaderboards and LMArena. Scores are refreshed regularly as new evaluations are published and new models are released.

No. While SWE-bench Verified is an important indicator, real-world performance depends on many factors including pricing, latency, context window, and specific task requirements. We recommend using our composite score which weighs multiple benchmarks and practical factors.

相关基准测试

SWE-bench Verified Benchmark - AI Code Generation Leaderboard (2026) | LM Market Cap