评分分布
300个AI模型综合评分分布的统计分析。探索均值、中位数、百分位数和层级分布,了解AI模型格局。
关键统计
所有300个评分模型的汇总统计。
平均分
68.2
+/- 18.3 标准差
中位分
71.9
评分范围
40-96
第95百分位
92.2
高于中位数
150
共300个模型
评分分布(10分区间)
评分分布
每个10分区间中的模型数量。
评分层级
按性能层级分组的模型及汇总统计。
| 层级 | 范围 | 数量 | 占比 |
|---|---|---|---|
| Elite | 90–100 | 30 | 10.0% |
| Strong | 70–89 | 135 | 45.0% |
| Average | 50–69 | 55 | 18.3% |
| Below Average | 30–49 | 73 | 24.3% |
| Weak | 0–29 | 0 | 0.0% |
百分位分析
关键百分位数的评分阈值。
| 百分位 | 评分 | 位置 |
|---|---|---|
| P5 | 40.0 | 4096 |
| P10 | 40.0 | 4096 |
| P25 | 53.2 | 4096 |
| P50 | 71.9 | 4096 |
| P75 | 81.8 | 4096 |
| P90 | 89.1 | 4096 |
| P95 | 92.2 | 4096 |
按平均分排名的服务商
拥有3+模型的服务商,按平均综合评分排名。
| 提供商 | 模型 | 平均评分 |
|---|---|---|
1xAI | 7 | 87.6 |
2Anthropic | 28 | 80.7 |
3OpenAI | 79 | 79.4 |
4Google | 28 | 75.6 |
5Zhipu AI | 18 | 74.4 |
6MiniMax | 7 | 71.1 |
7thinkingmachines | 4 | 70.8 |
8Alibaba | 31 | 66.9 |
9Mistral AI | 5 | 66.8 |
10Meta | 5 | 61.5 |
评分集中度
模型在前20%、中间60%和后20%评分中的分布情况。
方法论
评分的计算方式以及分布所揭示的信息。
评分计算方式
每个模型获得0到100的评分,主要基于基准测试分数(90%),来源包括Arena Elo、MMLU、GPQA、HumanEval、SWE-bench等15+标准化评估。功能和上下文窗口作为辅助排序(10%)。该评分旨在用一个数字衡量模型的科学评估质量。
分布图告诉我们什么
评分分布揭示了AI模型的竞争格局。中位数附近的紧密聚集表明有许多能力相近的模型,而分散的分布则表明层级之间有明显的差异。分布的形状、偏斜度以及均值和中位数之间的差距都能揭示市场是头重脚轻、底部沉重还是均匀分布。
The score distribution shows how all 290+ tracked AI models are spread across the 0-100 SignalScore scale. Most models cluster in the 40-70 range, with a small elite group scoring above 80 and budget/older models falling below 30.
SignalScore is a composite metric combining six weighted factors: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Each factor is normalized to a 0-100 scale before weighting.
Models scoring above the 75th percentile (typically 65+ SignalScore) are considered strong performers. The top 10% of models score above 78, while the median score across all tracked models sits around 52-55.