Skip to content

Ranking Confidence Explorer

Analyzes how confident we are in each model's ranking. Rank spreads show the range of positions a model could realistically hold, confidence levels indicate ranking precision, and stability states reflect consistency over time.

Confidence Level Distribution

LMMarketCap.com

Confidence Overview

High-level summary of ranking confidence across all 300 models.

High Confidence

300

100.0% of models

Medium Confidence

0

0.0% of models

Low Confidence

0

0.0% of models

Avg Rank Spread

4.0

positions of uncertainty

Confidence Distribution

Breakdown of models by confidence level with averages for score, spread, and rank.

Confidence LevelCount%
High300100.0%
Medium00.0%
Low00.0%

Most Precisely Ranked

Models with the tightest rank spreads. These are the rankings we are most confident about.

#ModelScoreRankSpread
1Claude Fable 597.11±2
2Claude Fable 5 (batch)97.12±3
3Claude Opus 5 (Fast)95.13±4
4Claude Opus 595.14±4
5Claude Opus 4.8 (Fast)95.15±4
6Claude Opus 4.895.16±4
7Claude Opus 4.7 (Fast)95.17±4
8Claude Opus 4.795.18±4
9Claude Opus 4.7 (batch)95.19±4
10Claude Opus 4.8 (batch)94.610±4
11GPT-5.5 Pro92.711±4
12GPT-5.5 Pro (batch)92.712±4
13GPT-5.592.713±4
14GPT-5.5 (batch)92.714±4
15Gemini 3.1 Pro Preview Custom Tools92.215±4
16Gemini 3.1 Pro Preview92.216±4
17Gemini 3.1 Pro Preview (batch)92.217±4
18GPT-5.4 Pro91.918±4
19GPT-5.4 Pro (batch)91.919±4
20GPT-5.491.920±4

Most Uncertain Rankings

Models with the widest rank spreads. These models could be ranked very differently under slight changes.

#ModelScoreRankSpread
1Claude Opus 5 (Fast)95.13±4
2Claude Opus 595.14±4
3Claude Opus 4.8 (Fast)95.15±4
4Claude Opus 4.895.16±4
5Claude Opus 4.7 (Fast)95.17±4
6Claude Opus 4.795.18±4
7Claude Opus 4.7 (batch)95.19±4
8Claude Opus 4.8 (batch)94.610±4
9GPT-5.5 Pro92.711±4
10GPT-5.5 Pro (batch)92.712±4
11GPT-5.592.713±4
12GPT-5.5 (batch)92.714±4
13Gemini 3.1 Pro Preview Custom Tools92.215±4
14Gemini 3.1 Pro Preview92.216±4
15Gemini 3.1 Pro Preview (batch)92.217±4
16GPT-5.4 Pro91.918±4
17GPT-5.4 Pro (batch)91.919±4
18GPT-5.491.920±4
19GPT-5.4 (batch)91.921±4
20GPT-5.3 Chat90.522±4

State × Confidence Matrix

Cross-tabulation of confidence level and stability state. The best combination is high confidence + stable; the worst is low confidence + fragile.

ConfidenceStableHeldFragilePreliminary
High13021968
Medium0000
Low0000

Rank Spread Visualization

Visual representation of ranking uncertainty for the top 30 models. The bar shows the possible rank range at 90% confidence; the marker shows the actual rank.

Rank 1Rank 32

What Affects Confidence

How ranking confidence is determined and what the metrics mean.

Rank Spread

Computed via bootstrap resampling of the scoring pipeline. By running thousands of simulations with slight variations, we determine the range of ranks each model could realistically hold. The spread represents the 90% confidence interval: in 9 out of 10 cases, the model's true rank falls within this range.

Confidence Level

Derived from the rank spread width. Models with tight spreads (small uncertainty) receive high confidence, meaning their ranking position is reliable. Wider spreads indicate medium or low confidence, where the model's position could shift significantly with different weighting or data updates.

State

A stability classification based on the consistency of performance metrics over time. “Stable” models show consistent rankings, “held” models maintain position with some variance, “fragile” models are prone to rank shifts, and “preliminary” models lack enough data history to assess stability.

Explore More

Continue exploring AI model data with other explorers and trackers.

Frequently Asked Questions

Ranking confidence is calculated using bootstrap resampling - a statistical technique that re-runs the ranking process thousands of times with slight variations to see how stable each model's position is. Models with narrow rank spreads have high confidence, while those with wide spreads have uncertain rankings.

Rank spread is the range between a model's best and worst possible rank across bootstrap simulations. A rank spread of 2 means the model might move 1 position up or down, while a spread of 20 means its true ranking is quite uncertain.

Low confidence usually means the model scores are clustered closely together with many competitors, making the exact ordering sensitive to small measurement differences. Models in the middle of the leaderboard tend to have wider rank spreads than those at the very top or bottom.

AI Ranking Confidence Explorer | LM Market Cap