Skip to content

排名置信度探索器

分析我们对每个模型排名的置信度。排名范围展示模型可能持有的位置范围,置信度水平表示排名精确度,稳定性状态反映随时间的一致性。

置信度水平分布

LMMarketCap.com

置信度概览

全部300个模型的排名置信度概览。

高置信度

300

100.0% of models

中等置信度

0

0.0% of models

低置信度

0

0.0% of models

平均排名跨度

4.0

位的不确定性

置信度分布

按置信度水平分类的模型细分,包括评分、范围和排名的平均值。

置信度级别数量%
High300100.0%
Medium00.0%
Low00.0%

排名最精确的

排名范围最窄的模型。这些是我们最有信心的排名。

#模型评分排名跨度
1Claude Fable 597.11±2
2Claude Fable 5 (batch)97.12±3
3Claude Opus 5 (Fast)95.13±4
4Claude Opus 595.14±4
5Claude Opus 4.8 (Fast)95.15±4
6Claude Opus 4.895.16±4
7Claude Opus 4.7 (Fast)95.17±4
8Claude Opus 4.795.18±4
9Claude Opus 4.7 (batch)95.19±4
10Claude Opus 4.8 (batch)94.610±4
11GPT-5.5 Pro92.711±4
12GPT-5.5 Pro (batch)92.712±4
13GPT-5.592.713±4
14GPT-5.5 (batch)92.714±4
15Gemini 3.1 Pro Preview Custom Tools92.215±4
16Gemini 3.1 Pro Preview92.216±4
17Gemini 3.1 Pro Preview (batch)92.217±4
18GPT-5.4 Pro91.918±4
19GPT-5.4 Pro (batch)91.919±4
20GPT-5.491.920±4

排名最不确定的

排名范围最宽的模型。这些模型在细微变化下可能排名差异很大。

#模型评分排名跨度
1Claude Opus 5 (Fast)95.13±4
2Claude Opus 595.14±4
3Claude Opus 4.8 (Fast)95.15±4
4Claude Opus 4.895.16±4
5Claude Opus 4.7 (Fast)95.17±4
6Claude Opus 4.795.18±4
7Claude Opus 4.7 (batch)95.19±4
8Claude Opus 4.8 (batch)94.610±4
9GPT-5.5 Pro92.711±4
10GPT-5.5 Pro (batch)92.712±4
11GPT-5.592.713±4
12GPT-5.5 (batch)92.714±4
13Gemini 3.1 Pro Preview Custom Tools92.215±4
14Gemini 3.1 Pro Preview92.216±4
15Gemini 3.1 Pro Preview (batch)92.217±4
16GPT-5.4 Pro91.918±4
17GPT-5.4 Pro (batch)91.919±4
18GPT-5.491.920±4
19GPT-5.4 (batch)91.921±4
20GPT-5.3 Chat90.522±4

状态 × 置信度矩阵

置信度水平与稳定性状态的交叉表。最佳组合是高置信度+稳定;最差是低置信度+脆弱。

置信度StableHeldFragilePreliminary
High13021968
Medium0000
Low0000

排名跨度可视化

前30个模型的排名不确定性可视化表示。条形显示90%置信度下的可能排名范围;标记显示实际排名。

排名 1排名 32

影响置信度的因素

排名置信度如何确定以及各指标的含义。

排名波动

通过评分管道的自助重采样计算。通过运行数千次带有微小变化的模拟,我们确定每个模型可能实际持有的排名范围。范围代表90%的置信区间:在十次中有九次,模型的真实排名在此范围内。

置信度级别

由排名范围宽度得出。范围较窄(不确定性小)的模型获得高置信度,意味着其排名位置是可靠的。较宽的范围表示中等或低置信度,模型的位置可能在不同权重或数据更新下发生显著变化。

状态

基于性能指标随时间一致性的稳定性分类。"稳定"模型显示一致的排名,"保持"模型在一定波动下维持位置,"脆弱"模型容易发生排名变化,"初步"模型缺乏足够的数据历史来评估稳定性。

探索更多

继续使用其他探索器和追踪器探索AI模型数据。

Frequently Asked Questions

Ranking confidence is calculated using bootstrap resampling - a statistical technique that re-runs the ranking process thousands of times with slight variations to see how stable each model's position is. Models with narrow rank spreads have high confidence, while those with wide spreads have uncertain rankings.

Rank spread is the range between a model's best and worst possible rank across bootstrap simulations. A rank spread of 2 means the model might move 1 position up or down, while a spread of 20 means its true ranking is quite uncertain.

Low confidence usually means the model scores are clustered closely together with many competitors, making the exact ordering sensitive to small measurement differences. Models in the middle of the leaderboard tend to have wider rank spreads than those at the very top or bottom.

AI Ranking Confidence Explorer | LM Market Cap