Skip to content

模型稳定性报告

哪些AI模型随时间最为一致?本报告分析了 300 个被追踪模型的排名变化、状态分类和波动曲线,生成0到100的稳定性评分。

坚如磐石

79

一致

11

可变

70

波动

140

稳定性分类分布

LMMarketCap.com

服务商稳定性排名 (平均分)

LMMarketCap.com

最稳定模型

稳定性评分最高的前20个模型。这些模型保持一致的排名,波动性最小。

#模型评分稳定性24小时7天
1Claude Fable 5Anthropic97.110000
2Claude Fable 5 (batch)Anthropic97.110000
3Claude Opus 4.8 (Fast)Anthropic95.11000+1
4Claude Opus 4.8Anthropic95.11000+1
5Claude Opus 4.8 (batch)Anthropic94.61000-2
6GPT-5.5 Pro (batch)OpenAI92.71000-3
7GPT-5.5 (batch)OpenAI92.71000-3
8Gemini 3.1 Pro Preview (batch)Google92.21000-3
9GPT-5.4 Pro (batch)OpenAI91.91000-3
10GPT-5.4 (batch)OpenAI91.91000-3
11Grok Build 0.1xAI40.01000-3
12Perceptron Mk1perceptron40.01000-3
13Ring-2.6-1Tinclusionai40.01000-3
14GPT Chat LatestOpenAI40.01000-3
15Nemotron 3 Nano Omni (free)NVIDIA40.01000-3
16Anthropic Claude Haiku Latest~anthropic40.01000-3
17OpenAI GPT Mini Latest~openai40.01000-3
18Google Gemini Pro Latest~google40.01000-3
19MoonshotAI Kimi Latest~moonshotai40.01000-3
20Google Gemini Flash Latest~google40.01000-3

最不稳定模型

稳定性评分最低的后20个模型。这些模型表现出显著的排名波动或不一致的状态。

#模型评分稳定性24小时7天
1GPT-5.1-Codex-MiniOpenAI49.819-167-174
2GLM 4.5Zhipu AI73.619-13-23
3GLM 5Zhipu AI82.122+29+18
4GLM 4.6Zhipu AI73.623-14+12
5Qwen3.6 PlusAlibaba74.124+4-7
6R1DeepSeek74.224+4-7
7o1OpenAI74.424+4-8
8Claude Sonnet 4Anthropic74.424+4-8
9Gemini 2.5 FlashGoogle79.124-4-12
10Gemini 2.5 Flash LiteGoogle79.124-4-12
11R1 0528DeepSeek79.424-4-12
12Muse Spark 1.2meta80.224-4+138
13DeepSeek V3.2DeepSeek81.324-4-12
14o4 MiniOpenAI81.424-4-12
15Claude Opus 4Anthropic82.124-4-12
16Claude Opus 4.1Anthropic82.125-4+40
17Qwen3.7 MaxAlibaba74.825+4+135
18Qwen3.5 397B A17BAlibaba79.426-4-12
19Qwen3.5 Plus 2026-04-20Alibaba74.827+4+147
20GLM 5 TurboZhipu AI82.127+28+67

各服务商稳定性

各服务商的汇总稳定性指标。服务商按所有模型的平均稳定性评分排名。

提供商模型平均稳定性
perceptron1100.0
~openai2100.0
~google2100.0
~moonshotai1100.0
Writer1100.0
Upstage199.7
~anthropic499.5
rekaai198.3
sakana198.0
Kuaishou396.7
aion-labs396.5
NVIDIA996.3
poolside495.0
Meituan195.0
~x-ai195.0
inclusionai591.8
ByteDance485.5
StepFun269.7
xAI567.4
Anthropic2864.8
~deepseek164.0
TII364.0
OpenAI8658.8
DeepSeek1255.4
Cohere455.0
Moonshot AI855.0
Amazon154.1
Google2753.6
Inception150.3
Alibaba3149.6
Mistral AI648.6
Tencent248.5
MiniMax847.8
IBM146.8
thinkingmachines346.7
Allen AI143.1
Xiaomi242.5
arcee-ai141.7
Meta540.5
Cursor240.0
Microsoft139.0
Zhipu AI1338.3
meta233.6

稳定性分布

所有 300 个被追踪模型的稳定性评分分布。

0–10
0
10–20
2
20–30
18
30–40
55
40–50
65
50–60
35
60–70
35
70–80
3
80–90
16
90–100
71

什么让模型保持稳定?

我们的稳定性评分系统使用三个关键信号来衡量模型随时间的一致性表现。

排名一致性

稳定性的最直接衡量标准。模型因24小时内较大的排名变化最多失去25分(每移动一个排名位置扣5分),7天变化最多失去21分(每个位置扣3分)。排名保持稳定的模型得分更高。

状态分类

每个模型都有一个反映其整体可靠性的状态。处于"稳定"状态的模型获得10分加分,而"脆弱"模型被扣15分。这捕捉了超越简单排名变动的系统性可靠性。

波动曲线

14天的波动曲线数据揭示了隐藏的波动性。我们计算波动曲线的标准差并最多减去20分。即使最终回到起点的模型,如果中间剧烈波动也会被扣分。

相关

Frequently Asked Questions

The stability score starts at 100 and is reduced based on three factors: 24-hour rank changes (up to -25 points, at 5 per position moved), 7-day rank changes (up to -21 points, at 3 per position), and sparkline volatility measured by standard deviation (up to -20 points). Models in a "stable" state get a +10 bonus, while "fragile" models lose 15 points.

Models are classified into four tiers based on their stability score: "Rock Solid" (85-100) means extremely consistent performance with minimal fluctuation. "Consistent" (70-84) means generally reliable with minor variations. "Variable" (50-69) shows noticeable ranking fluctuations. "Volatile" (below 50) indicates significant instability and unpredictable performance.

Stability indicates how predictably a model will perform over time. A highly rated but volatile model may deliver inconsistent results, which is problematic for production applications requiring reliable output quality. Stable models provide more predictable performance, making them safer choices for mission-critical workloads even if they do not always hold the top rank.

AI Model Stability Report — Consistency Rankings | LM Market Cap