Skip to content

历史表现 - 编程

顶级编程AI模型的90天评分趋势。追踪综合评分变化、模型入榜日期以及影响排名的关键事件。

评分趋势

排名模型提供商评分趋势30d60d90d
#1Claude Fable 5Anthropic97.1+0.5+0.5+0.5
#2Claude Fable 5 (batch)Anthropic97.10.00.00.0
#3Claude Opus 5 (Fast)Anthropic95.1+55.1+55.1+55.1
#4Claude Opus 5Anthropic95.1+55.1+55.1+55.1
#5Claude Opus 4.8 (Fast)Anthropic95.1+12.1+12.1+12.1
#6Claude Opus 4.8Anthropic95.1+12.1+12.1+12.1
#7Claude Opus 4.7 (Fast)Anthropic95.1+0.4+0.4+0.4
#8Claude Opus 4.7Anthropic95.1+13.6+13.6+13.6
#9Claude Opus 4.7 (batch)Anthropic95.10.00.00.0
#10Claude Opus 4.8 (batch)Anthropic94.60.00.00.0

模型时间线

进入日期模型提供商进入排名当前排名
2025-11-07Claude Opus 4.7 (batch)Anthropic#12#9
2025-10-29Claude Opus 4.8 (batch)Anthropic#13#10
2025-10-05Claude Fable 5Anthropic#4#1
2025-10-02Claude Opus 4.8Anthropic#9#6
2025-09-30Claude Opus 5Anthropic#11#4
2025-08-20Claude Opus 5 (Fast)Anthropic#9#3
2025-08-02Claude Opus 4.7Anthropic#11#8
2025-07-28Claude Opus 4.8 (Fast)Anthropic#8#5
2025-07-11Claude Opus 4.7 (Fast)Anthropic#10#7
2025-06-02Claude Fable 5 (batch)Anthropic#5#2

关键事件

2026-02-28versionClaude Opus 4.6 released with expanded context
2026-02-15pricingOpenAI reduced GPT-5.2 pricing by 20%
2026-02-01versionGemini 3 Pro launched with multimodal improvements
2026-01-20versionDeepSeek V3.1 update with enhanced reasoning
2026-01-10pricingAnthropic introduced new Claude Sonnet tier pricing
2025-12-15versionQwen 3.5 397B released by Alibaba Cloud
2025-12-01pricingGoogle adjusted Gemini API pricing structure
2025-11-15versionGrok 4.1 launched with code generation focus

相关

Frequently Asked Questions

We track composite scores for the top coding AI models over a 90-day rolling window. Scores combine coding benchmarks like SWE-bench and HumanEval, pricing, context window, and capability data that refreshes hourly.

These columns show how much each model's composite score has changed over the last 30, 60, or 90 days. A positive change indicates improving performance or rankings, while a negative change suggests the model is falling behind newer competitors.

Check the Score Trends table above to see which models show the largest positive 30-day change. New model releases and major updates often cause significant score improvements.

Score data is refreshed hourly. The historical trend lines and change percentages are recalculated with each update to reflect the latest benchmark results, pricing changes, and capability additions.

Coding AI Historical Performance - Score Trends Over Time | LM Market Cap