Skip to content

Score Distribution

Statistical analysis of how 300 AI model composite scores are distributed. Explore the mean, median, percentiles, and tier breakdowns to understand the AI model landscape.

Key Statistics

Summary statistics across all 300 scored models.

Mean Score

67.8

+/- 18.5 stddev

Median Score

71.8

Score Range

40-97

95th Percentile

92.2

Above Median

152

of 300 models

Score Distribution (10-point buckets)

LMMarketCap.com

Score Distribution

Number of models in each 10-point score bucket.

0-10
0
10-20
0
20-30
0
30-40
0
40-50
74
50-60
18
60-70
48
70-80
72
80-90
57
90-100
31
Elite
Strong
Average
Below Average
Weak

Score Tiers

Models grouped by performance tier with summary statistics.

TierRangeCount% of Total
Elite901003110.3%
Strong708912943.0%
Average50695819.3%
Below Average30497424.7%
Weak02900.0%

Percentile Breakdown

Score thresholds at key percentile levels.

PercentileScorePosition
P540.0
4097
P1040.0
4097
P2551.2
4097
P5071.8
4097
P7582.2
4097
P9090.4
4097
P9592.2
4097

Top Providers by Average Score

Providers with 3+ models, ranked by average composite score.

ProviderModelsAvg Score
1Anthropic
2882.5
2xAI
578.9
3Google
2778.5
4OpenAI
8677.6
5MiniMax
873.5
6Zhipu AI
1373.1
7thinkingmachines
372.8
8DeepSeek
1266.4
9Alibaba
3165.0
10Mistral AI
662.3

Score Concentration

How models are distributed across the top 20%, middle 60%, and bottom 20% of scores.

Top 20%(Score >= 86.7)
64 (21.3%)
Middle 60%(Score 40.0 - 86.7)
169 (56.3%)
Bottom 20%(Score <= 40.0)
67 (22.3%)

Methodology

How scores are computed and what the distribution reveals.

How Scores Are Computed

Each model receives a composite score from 0 to 100, calculated as a weighted combination of six signals: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). The score is designed to capture overall model quality and value in a single number.

What the Distribution Tells Us

The score distribution reveals the competitive landscape of AI models. A tight cluster near the median suggests many similarly capable models, while a wide spread indicates clear differentiation between tiers. The shape of the distribution, its skew, and the gap between mean and median all provide insight into whether the market is top-heavy, bottom-heavy, or evenly distributed.

Explore More

Continue exploring AI model data with benchmarks, capabilities, and the full leaderboard.

Frequently Asked Questions

The score distribution shows how all 290+ tracked AI models are spread across the 0-100 SignalScore scale. Most models cluster in the 40-70 range, with a small elite group scoring above 80 and budget/older models falling below 30.

SignalScore is a composite metric combining six weighted factors: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Each factor is normalized to a 0-100 scale before weighting.

Models scoring above the 75th percentile (typically 65+ SignalScore) are considered strong performers. The top 10% of models score above 78, while the median score across all tracked models sits around 52-55.

AI Model Score Distribution - Statistical Analysis | LM Market Cap