Skip to content

模型稳定性报告

哪些AI模型随时间最为一致?本报告分析了 300 个被追踪模型的排名变化、状态分类和波动曲线,生成0到100的稳定性评分。

坚如磐石

36

一致

18

可变

84

波动

162

稳定性分类分布

LMMarketCap.com

服务商稳定性排名 (平均分)

LMMarketCap.com

最稳定模型

稳定性评分最高的前20个模型。这些模型保持一致的排名,波动性最小。

#模型评分稳定性24小时7天
1Claude Fable 5.1Anthropic95.910000
2Claude Fable 5.1 (batch)Anthropic95.910000
3Claude Fable 5Anthropic95.910000
4Claude Fable 5 (batch)Anthropic95.910000
5Claude Opus 4.8Anthropic95.110000
6Claude Opus 4.7 (batch)Anthropic95.11000-3
7Claude Opus 4.8 (batch)Anthropic94.61000-2
8GPT-5.5 Pro (batch)OpenAI92.71000-3
9GPT-5.5 (batch)OpenAI92.71000-3
10Gemini 3.1 Pro Preview (batch)Google92.21000-3
11GPT-5.4 Pro (batch)OpenAI91.91000-3
12GPT-5.4 (batch)OpenAI91.91000-3
13Grok 4.6xAI88.81000+3
14GPT-5.2 Pro (batch)OpenAI90.5980-4
15GPT-5.2 (batch)OpenAI90.5980-4
16GPT-5.6 Luna Pro (batch)OpenAI89.0980-4
17GPT-5.6 Luna (batch)OpenAI89.0980-4
18GPT-5.6 Terra Pro (batch)OpenAI89.0980-4
19GPT-5.6 Terra (batch)OpenAI89.0980-4
20GPT-5.6 Sol Pro (batch)OpenAI89.0980-4

最不稳定模型

稳定性评分最低的后20个模型。这些模型表现出显著的排名波动或不一致的状态。

#模型评分稳定性24小时7天
1Command R+ (08-2024)Cohere48.734-2-11
2Command ACohere50.834-2-11
3Claude 3 HaikuAnthropic51.334-2-11
4GPT-4o-mini (2024-07-18)OpenAI56.534-2-12
5Phi 4Microsoft60.234-2-11
6Qwen3 8BAlibaba61.034-2-11
7Llama 4 ScoutMeta60.235-2-11
8Qwen3 235B A22BAlibaba54.036-2-11
9Llama 3.1 8B InstructMeta44.537-2-11
10gpt-oss-20bOpenAI57.438-2-11
11R1 Distill Llama 70BDeepSeek40.739-2-11
12Laguna S 2.1poolside40.039-6-19
13Ling 3.0 Flashinclusionai40.039-6-19
14Claude Opus 5 (batch)Anthropic40.039-6-19
15Qwen3.7 FlashAlibaba40.039-6-20
16DeepSeek V4 Flash 0731DeepSeek40.039-6-21
17DeepSeek V4 Flash Latest~deepseek40.039-6-21
18Muse Glimmer 30Bmeta40.039-6-23
19Solar Pro 4Upstage40.039-6-23
20Sakana Namazusakana40.039-6-23

各服务商稳定性

各服务商的汇总稳定性指标。服务商按所有模型的平均稳定性评分排名。

提供商模型平均稳定性
aion-labs279.0
xAI773.9
Anthropic2768.8
OpenAI8863.7
Upstage259.0
thinkingmachines457.6
Google2657.0
Xiaomi554.3
fireworks154.0
IBM153.8
Moonshot AI252.9
Zhipu AI1952.6
Inception249.9
Amazon149.5
MiniMax748.3
Alibaba3346.5
meta644.7
Tencent644.7
DeepSeek1443.6
arcee-ai143.5
Mistral AI543.0
StepFun142.7
Cursor241.9
prism-ml139.0
unbiased139.0
~deepseek339.0
inference-net239.0
~openai439.0
sakana339.0
inclusionai539.0
~z-ai239.0
dots-studio139.0
ByteDance239.0
Liquid AI139.0
NVIDIA239.0
poolside139.0
Meta537.8
Cohere437.8
Microsoft134.0

稳定性分布

所有 300 个被追踪模型的稳定性评分分布。

0–10
0
10–20
0
20–30
0
30–40
99
40–50
63
50–60
63
60–70
21
70–80
11
80–90
11
90–100
32

什么让模型保持稳定?

我们的稳定性评分系统使用三个关键信号来衡量模型随时间的一致性表现。

排名一致性

稳定性的最直接衡量标准。模型因24小时内较大的排名变化最多失去25分(每移动一个排名位置扣5分),7天变化最多失去21分(每个位置扣3分)。排名保持稳定的模型得分更高。

状态分类

每个模型都有一个反映其整体可靠性的状态。处于"稳定"状态的模型获得10分加分,而"脆弱"模型被扣15分。这捕捉了超越简单排名变动的系统性可靠性。

波动曲线

14天的波动曲线数据揭示了隐藏的波动性。我们计算波动曲线的标准差并最多减去20分。即使最终回到起点的模型,如果中间剧烈波动也会被扣分。

相关

Frequently Asked Questions

The stability score starts at 100 and is reduced based on three factors: 24-hour rank changes (up to -25 points, at 5 per position moved), 7-day rank changes (up to -21 points, at 3 per position), and sparkline volatility measured by standard deviation (up to -20 points). Models in a "stable" state get a +10 bonus, while "fragile" models lose 15 points.

Models are classified into four tiers based on their stability score: "Rock Solid" (85-100) means extremely consistent performance with minimal fluctuation. "Consistent" (70-84) means generally reliable with minor variations. "Variable" (50-69) shows noticeable ranking fluctuations. "Volatile" (below 50) indicates significant instability and unpredictable performance.

Stability indicates how predictably a model will perform over time. A highly rated but volatile model may deliver inconsistent results, which is problematic for production applications requiring reliable output quality. Stable models provide more predictable performance, making them safer choices for mission-critical workloads even if they do not always hold the top rank.

AI Model Stability Report — Consistency Rankings | LM Market Cap