Score Distribution
Statistical analysis of how 300 AI model composite scores are distributed. Explore the mean, median, percentiles, and tier breakdowns to understand the AI model landscape.
Key Statistics
Summary statistics across all 300 scored models.
Mean Score
67.8
+/- 18.5 stddev
Median Score
71.8
Score Range
40-97
95th Percentile
92.2
Above Median
152
of 300 models
Score Distribution (10-point buckets)
Score Distribution
Number of models in each 10-point score bucket.
Score Tiers
Models grouped by performance tier with summary statistics.
| Tier | Range | Count | % of Total |
|---|---|---|---|
| Elite | 90–100 | 31 | 10.3% |
| Strong | 70–89 | 129 | 43.0% |
| Average | 50–69 | 58 | 19.3% |
| Below Average | 30–49 | 74 | 24.7% |
| Weak | 0–29 | 0 | 0.0% |
Percentile Breakdown
Score thresholds at key percentile levels.
| Percentile | Score | Position |
|---|---|---|
| P5 | 40.0 | 4097 |
| P10 | 40.0 | 4097 |
| P25 | 51.2 | 4097 |
| P50 | 71.8 | 4097 |
| P75 | 82.2 | 4097 |
| P90 | 90.4 | 4097 |
| P95 | 92.2 | 4097 |
Top Providers by Average Score
Providers with 3+ models, ranked by average composite score.
| Provider | Models | Avg Score |
|---|---|---|
1Anthropic | 28 | 82.5 |
2xAI | 5 | 78.9 |
3Google | 27 | 78.5 |
4OpenAI | 86 | 77.6 |
5MiniMax | 8 | 73.5 |
6Zhipu AI | 13 | 73.1 |
7thinkingmachines | 3 | 72.8 |
8DeepSeek | 12 | 66.4 |
9Alibaba | 31 | 65.0 |
10Mistral AI | 6 | 62.3 |
Score Concentration
How models are distributed across the top 20%, middle 60%, and bottom 20% of scores.
Methodology
How scores are computed and what the distribution reveals.
How Scores Are Computed
Each model receives a composite score from 0 to 100, calculated as a weighted combination of six signals: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). The score is designed to capture overall model quality and value in a single number.
What the Distribution Tells Us
The score distribution reveals the competitive landscape of AI models. A tight cluster near the median suggests many similarly capable models, while a wide spread indicates clear differentiation between tiers. The shape of the distribution, its skew, and the gap between mean and median all provide insight into whether the market is top-heavy, bottom-heavy, or evenly distributed.
Explore More
Continue exploring AI model data with benchmarks, capabilities, and the full leaderboard.
The score distribution shows how all 290+ tracked AI models are spread across the 0-100 SignalScore scale. Most models cluster in the 40-70 range, with a small elite group scoring above 80 and budget/older models falling below 30.
SignalScore is a composite metric combining six weighted factors: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Each factor is normalized to a 0-100 scale before weighting.
Models scoring above the 75th percentile (typically 65+ SignalScore) are considered strong performers. The top 10% of models score above 78, while the median score across all tracked models sits around 52-55.