Software Engineering Benchmark (Verified) 排行榜
Can a model resolve real GitHub issues from popular Python repositories? Human-validated subset ensures accurate evaluation. Tests end-to-end software engineering ability.
为什么重要: The gold standard for real-world coding ability. Unlike HumanEval, tests understanding of large codebases, debugging, and complex changes. Scores range 20-80%.
顶级模型
95%
Claude Fable 5
平均评分
62.0%
共52个模型
已测试模型
52
指标: resolve rate
人类基准
-
评分范围: 0%–100%
SWE-bench Verified Scores - Top 25 Models
Ranked by SWE-bench Verified score (%)
模型排名
All models with a reported SWE-bench Verified score, ranked by highest resolve rate.
关于 SWE-bench Verified
- 全名
- Software Engineering Benchmark (Verified)
- 类别
- Coding
- 指标
- resolve rate (%)
- 评分范围
- 0%–100%
- 人类基准
- 尚未确定
- 状态
- 启用
SWE-bench Verified is a standardized evaluation that measures AI model performance on specific tasks. It provides comparable scores across different models, helping developers choose the right model for their needs.
Claude Fable 5 currently holds the top score on the SWE-bench Verified benchmark. See our full rankings table above for the complete leaderboard with 52 models.
We update benchmark data from multiple sources including HuggingFace open-source model leaderboards and LMArena. Scores are refreshed regularly as new evaluations are published and new models are released.
No. While SWE-bench Verified is an important indicator, real-world performance depends on many factors including pricing, latency, context window, and specific task requirements. We recommend using our composite score which weighs multiple benchmarks and practical factors.