Skip to content

Score Distribution

Statistical analysis of how 300 AI model composite scores are distributed. Explore the mean, median, percentiles, and tier breakdowns to understand the AI model landscape.

Key Statistics

Summary statistics across all 300 scored models.

Mean Score

68.2

+/- 18.3 stddev

Median Score

71.9

Score Range

40-96

95th Percentile

92.2

Above Median

150

of 300 models

Score Distribution (10-point buckets)

LMMarketCap.com

Score Distribution

Number of models in each 10-point score bucket.

0-10
0
10-20
0
20-30
0
30-40
0
40-50
73
50-60
11
60-70
51
70-80
78
80-90
57
90-100
30
Elite
Strong
Average
Below Average
Weak

Score Tiers

Models grouped by performance tier with summary statistics.

TierRangeCount% of Total
Elite901003010.0%
Strong708913545.0%
Average50695518.3%
Below Average30497324.3%
Weak02900.0%

Percentile Breakdown

Score thresholds at key percentile levels.

PercentileScorePosition
P540.0
4096
P1040.0
4096
P2553.2
4096
P5071.9
4096
P7581.8
4096
P9089.1
4096
P9592.2
4096

Top Providers by Average Score

Providers with 3+ models, ranked by average composite score.

ProviderModelsAvg Score
1xAI
787.6
2Anthropic
2782.2
3OpenAI
8775.8
4Google
2875.6
5Zhipu AI
1874.4
6MiniMax
771.1
7thinkingmachines
470.8
8Mistral AI
566.8
9Alibaba
3266.1
10Meta
561.5

Score Concentration

How models are distributed across the top 20%, middle 60%, and bottom 20% of scores.

Top 20%(Score >= 86.7)
62 (20.7%)
Middle 60%(Score 40.0 - 86.7)
171 (57.0%)
Bottom 20%(Score <= 40.0)
67 (22.3%)

Methodology

How scores are computed and what the distribution reveals.

How Scores Are Computed

Each model receives a composite score from 0 to 100, calculated as a weighted combination of six signals: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). The score is designed to capture overall model quality and value in a single number.

What the Distribution Tells Us

The score distribution reveals the competitive landscape of AI models. A tight cluster near the median suggests many similarly capable models, while a wide spread indicates clear differentiation between tiers. The shape of the distribution, its skew, and the gap between mean and median all provide insight into whether the market is top-heavy, bottom-heavy, or evenly distributed.

Explore More

Continue exploring AI model data with benchmarks, capabilities, and the full leaderboard.

Frequently Asked Questions

The score distribution shows how all 290+ tracked AI models are spread across the 0-100 SignalScore scale. Most models cluster in the 40-70 range, with a small elite group scoring above 80 and budget/older models falling below 30.

SignalScore is a composite metric combining six weighted factors: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Each factor is normalized to a 0-100 scale before weighting.

Models scoring above the 75th percentile (typically 65+ SignalScore) are considered strong performers. The top 10% of models score above 78, while the median score across all tracked models sits around 52-55.

AI Model Score Distribution - Statistical Analysis | LM Market Cap