Skip to content

Model Stability Report

Which AI models are the most consistent over time? This report analyzes rank changes, state classifications, and sparkline volatility across 300 tracked models to produce a stability score from 0 to 100.

Rock Solid

10

Consistent

71

Variable

102

Volatile

117

Stability Classification Distribution

LMMarketCap.com

Provider Stability Rankings (Avg Score)

LMMarketCap.com

Most Stable Models

Top 20 models with the highest stability scores. These models maintain consistent rankings with minimal volatility.

#ModelScoreStability24h7d
1Claude Fable 5Anthropic97.110000
2Claude Opus 4.8 (Fast)Anthropic95.11000-1
3Claude Opus 4.8Anthropic95.11000-1
4GPT-5.5 ProOpenAI92.710000
5Falcon-H1-Arabic 34B InstructTII40.01000+1
6Falcon-H1-Arabic 7B InstructTII40.01000+1
7Falcon-H1-Arabic 3B InstructTII40.01000+1
8MiniMax M2.7MiniMax78.0950-1
9Claude Opus 4.7 (Fast)Anthropic95.1950-5
10GPT-5.3-CodexOpenAI90.59300
11Claude Opus 4.1Anthropic82.1850-2
12Claude Opus 4.7Anthropic95.1830-5
13Claude Fable 5 (batch)Anthropic97.1790+377
14Claude Opus 4.7 (batch)Anthropic95.1790+370
15Claude Opus 4.8 (batch)Anthropic94.6790+369
16GPT-5.5 Pro (batch)OpenAI92.7790+367
17GPT-5.5 (batch)OpenAI92.7790+365
18Gemini 3.1 Pro Preview (batch)Google92.2790+362
19GPT-5.4 Pro (batch)OpenAI91.9790+360
20GPT-5.4 (batch)OpenAI91.9790+358

Most Volatile Models

Bottom 20 models with the lowest stability scores. These models show significant ranking fluctuations or inconsistent states.

#ModelScoreStability24h7d
1Command R (08-2024)Cohere48.739+1-62
2Kimi K3Moonshot AI40.0440-63
3Command R+ (08-2024)Cohere48.7440-63
4Command ACohere50.8440-63
5Claude 3 HaikuAnthropic51.3440-63
6Kimi K2 0711Moonshot AI51.4440-63
7GPT-4o-mini (2024-07-18)OpenAI56.5440-64
8gpt-oss-20b (free)OpenAI57.4440-64
9Phi 4Microsoft60.2440-63
10Qwen3 8BAlibaba61.0440-63
11Mixtral 8x22B InstructMistral AI63.4440-61
12o3 Mini HighOpenAI64.1440-63
13Qwen3 30B A3BAlibaba64.1440-63
14Qwen3 30B A3B Thinking 2507Alibaba64.1440-59
15Qwen3 235B A22B Instruct 2507Alibaba64.7440-61
16GPT-4OpenAI64.8440-61
17GPT-4 Turbo PreviewOpenAI64.8440-45
18Llama 3.1 70B InstructMeta65.3440-60
19Qwen3 235B A22B Thinking 2507Alibaba65.5440-60
20Mistral LargeMistral AI65.9440-60

Stability by Provider

Aggregated stability metrics per provider. Providers are ranked by their average stability score across all models.

ProviderModelsAvg Stability
TII3100.0
xAI579.0
meta271.1
inclusionai570.0
Anthropic2869.6
thinkingmachines365.0
~deepseek164.0
poolside464.0
Meituan164.0
~x-ai164.0
sakana164.0
~anthropic464.0
perceptron164.0
~openai264.0
~google264.0
~moonshotai164.0
OpenAI8662.9
Kuaishou362.8
aion-labs362.5
NVIDIA962.4
Tencent261.0
Google2760.4
Amazon159.1
Writer158.5
rekaai158.2
Moonshot AI857.3
MiniMax857.2
Upstage156.6
Inception155.2
StepFun255.2
Zhipu AI1353.6
IBM151.8
Alibaba3151.2
DeepSeek1249.9
ByteDance448.2
Allen AI147.9
Cohere447.8
Xiaomi247.3
arcee-ai146.3
Mistral AI645.2
Cursor244.6
Meta544.4
Microsoft144.0

Stability Distribution

How stability scores are distributed across all 300 tracked models.

0–10
0
10–20
0
20–30
0
30–40
1
40–50
116
50–60
39
60–70
63
70–80
69
80–90
2
90–100
10

What Makes a Model Stable?

Our stability scoring system uses three key signals to measure how consistently a model performs over time.

Rank Consistency

The most direct measure of stability. Models lose up to 25 points for large 24-hour rank changes (5 points per rank position moved) and up to 21 points for 7-day changes (3 points per position). Models that hold their rank tightly score higher.

State Classification

Each model has a state reflecting its overall reliability. Models in a "stable" state receive a 10-point bonus, while "fragile" models are penalized 15 points. This captures systemic reliability beyond simple rank movement.

Sparkline Volatility

The 14-day sparkline data reveals hidden volatility. We compute the standard deviation of the sparkline and subtract up to 20 points. Even models that end where they started can be penalized if they oscillated wildly along the way.

Related

Frequently Asked Questions

The stability score starts at 100 and is reduced based on three factors: 24-hour rank changes (up to -25 points, at 5 per position moved), 7-day rank changes (up to -21 points, at 3 per position), and sparkline volatility measured by standard deviation (up to -20 points). Models in a "stable" state get a +10 bonus, while "fragile" models lose 15 points.

Models are classified into four tiers based on their stability score: "Rock Solid" (85-100) means extremely consistent performance with minimal fluctuation. "Consistent" (70-84) means generally reliable with minor variations. "Variable" (50-69) shows noticeable ranking fluctuations. "Volatile" (below 50) indicates significant instability and unpredictable performance.

Stability indicates how predictably a model will perform over time. A highly rated but volatile model may deliver inconsistent results, which is problematic for production applications requiring reliable output quality. Stable models provide more predictable performance, making them safer choices for mission-critical workloads even if they do not always hold the top rank.

AI Model Stability Report — Consistency Rankings | LM Market Cap