评分分布
300个AI模型综合评分分布的统计分析。探索均值、中位数、百分位数和层级分布,了解AI模型格局。
关键统计
所有300个评分模型的汇总统计。
平均分
67.8
+/- 18.5 标准差
中位分
71.8
评分范围
40-97
第95百分位
92.2
高于中位数
152
共300个模型
评分分布(10分区间)
评分分布
每个10分区间中的模型数量。
评分层级
按性能层级分组的模型及汇总统计。
| 层级 | 范围 | 数量 | 占比 |
|---|---|---|---|
| Elite | 90–100 | 31 | 10.3% |
| Strong | 70–89 | 129 | 43.0% |
| Average | 50–69 | 58 | 19.3% |
| Below Average | 30–49 | 74 | 24.7% |
| Weak | 0–29 | 0 | 0.0% |
百分位分析
关键百分位数的评分阈值。
| 百分位 | 评分 | 位置 |
|---|---|---|
| P5 | 40.0 | 4097 |
| P10 | 40.0 | 4097 |
| P25 | 51.2 | 4097 |
| P50 | 71.8 | 4097 |
| P75 | 82.2 | 4097 |
| P90 | 90.4 | 4097 |
| P95 | 92.2 | 4097 |
按平均分排名的服务商
拥有3+模型的服务商,按平均综合评分排名。
| 提供商 | 模型 | 平均评分 |
|---|---|---|
1Anthropic | 28 | 82.5 |
2xAI | 5 | 78.9 |
3Google | 27 | 78.5 |
4OpenAI | 86 | 77.6 |
5MiniMax | 8 | 73.5 |
6Zhipu AI | 13 | 73.1 |
7thinkingmachines | 3 | 72.8 |
8DeepSeek | 12 | 66.4 |
9Alibaba | 31 | 65.0 |
10Mistral AI | 6 | 62.3 |
评分集中度
模型在前20%、中间60%和后20%评分中的分布情况。
方法论
评分的计算方式以及分布所揭示的信息。
评分计算方式
每个模型获得0到100的评分,主要基于基准测试分数(90%),来源包括Arena Elo、MMLU、GPQA、HumanEval、SWE-bench等15+标准化评估。功能和上下文窗口作为辅助排序(10%)。该评分旨在用一个数字衡量模型的科学评估质量。
分布图告诉我们什么
评分分布揭示了AI模型的竞争格局。中位数附近的紧密聚集表明有许多能力相近的模型,而分散的分布则表明层级之间有明显的差异。分布的形状、偏斜度以及均值和中位数之间的差距都能揭示市场是头重脚轻、底部沉重还是均匀分布。
The score distribution shows how all 290+ tracked AI models are spread across the 0-100 SignalScore scale. Most models cluster in the 40-70 range, with a small elite group scoring above 80 and budget/older models falling below 30.
SignalScore is a composite metric combining six weighted factors: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Each factor is normalized to a 0-100 scale before weighting.
Models scoring above the 75th percentile (typically 65+ SignalScore) are considered strong performers. The top 10% of models score above 78, while the median score across all tracked models sits around 52-55.