Skip to content

Historical Performance - Coding

90-day score trends for top coding AI models. Track composite score changes, model entry dates, and key events that impacted rankings.

Score Trends

RankModelProviderScoreTrend30d60d90d
#1Claude Fable 5Anthropic97.1+0.5+0.5+0.5
#2Claude Fable 5 (batch)Anthropic97.10.00.00.0
#3Claude Opus 5 (Fast)Anthropic95.1+55.1+55.1+55.1
#4Claude Opus 5Anthropic95.1+55.1+55.1+55.1
#5Claude Opus 4.8 (Fast)Anthropic95.1+12.1+12.1+12.1
#6Claude Opus 4.8Anthropic95.1+12.1+12.1+12.1
#7Claude Opus 4.7 (Fast)Anthropic95.1+0.4+0.4+0.4
#8Claude Opus 4.7Anthropic95.1+13.6+13.6+13.6
#9Claude Opus 4.7 (batch)Anthropic95.10.00.00.0
#10Claude Opus 4.8 (batch)Anthropic94.60.00.00.0

Model Timeline

Date EnteredModelProviderEntry RankCurrent Rank
2025-11-07Claude Opus 4.7 (batch)Anthropic#12#9
2025-10-29Claude Opus 4.8 (batch)Anthropic#13#10
2025-10-05Claude Fable 5Anthropic#4#1
2025-10-02Claude Opus 4.8Anthropic#9#6
2025-09-30Claude Opus 5Anthropic#11#4
2025-08-20Claude Opus 5 (Fast)Anthropic#9#3
2025-08-02Claude Opus 4.7Anthropic#11#8
2025-07-28Claude Opus 4.8 (Fast)Anthropic#8#5
2025-07-11Claude Opus 4.7 (Fast)Anthropic#10#7
2025-06-02Claude Fable 5 (batch)Anthropic#5#2

Key Events

2026-02-28versionClaude Opus 4.6 released with expanded context
2026-02-15pricingOpenAI reduced GPT-5.2 pricing by 20%
2026-02-01versionGemini 3 Pro launched with multimodal improvements
2026-01-20versionDeepSeek V3.1 update with enhanced reasoning
2026-01-10pricingAnthropic introduced new Claude Sonnet tier pricing
2025-12-15versionQwen 3.5 397B released by Alibaba Cloud
2025-12-01pricingGoogle adjusted Gemini API pricing structure
2025-11-15versionGrok 4.1 launched with code generation focus

Related

Frequently Asked Questions

We track composite scores for the top coding AI models over a 90-day rolling window. Scores combine coding benchmarks like SWE-bench and HumanEval, pricing, context window, and capability data that refreshes hourly.

These columns show how much each model's composite score has changed over the last 30, 60, or 90 days. A positive change indicates improving performance or rankings, while a negative change suggests the model is falling behind newer competitors.

Check the Score Trends table above to see which models show the largest positive 30-day change. New model releases and major updates often cause significant score improvements.

Score data is refreshed hourly. The historical trend lines and change percentages are recalculated with each update to reflect the latest benchmark results, pricing changes, and capability additions.

Coding AI Historical Performance - Score Trends Over Time | LM Market Cap