Skip to content
Benchmark category

Arena Benchmarks

Compare top models across the benchmark suite that best represents arena performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.

2

Benchmarks in category

137

Models with coverage

0

Benchmarks with human baseline

0

Saturated benchmarks

What Is Included

The current benchmark set in this category, with context on what each test captures.

All benchmarks

LMSYS Chatbot Arena Elo Rating

The most trusted 'vibes-based' benchmark — reflects real human preferences, not just academic metrics. Widely considered the most meaningful overall ranking.

Elo rating

LiveBench (Dynamic)

Contamination-free by design — uses new questions regularly. Top models still score below 70%, making it highly discriminating.

average score %
Model type:

LMSYS Chatbot Arena Elo Rating

Human preference rating from 6M+ crowdsourced blind head-to-head comparisons. Users chat with two anonymous models and pick the better response.

Metric: Elo rating

Why it matters

The most trusted 'vibes-based' benchmark — reflects real human preferences, not just academic metrics. Widely considered the most meaningful overall ranking.

Arena Elo Scores (131 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇Claude Fable 51507
2🥈Claude Opus 4.61503
3🥉Qwen3.8 Max1497
4Gemini 3.1 Pro Preview1494
5Gemini 3.1 Pro1494
6Claude Opus 4.71491
7Muse Spark 1.11487
8Gemini 3 Pro1486
9GPT-5.41485
10Gemini 3.6 Flash1485
11GPT-5.21481
12Claude Opus 4.81479
13Gemini 3.5 Flash1477
14GPT-5.2 Chat1476
15GPT-5.11475
16GPT-5.51475
17Gemini 3 Flash1474
18Grok 4.51468
19GLM 5.11468
20MiMo-V2.5-Pro1467
21GPT-51465
22DeepSeek V4 Pro1463
23Grok 41462
24Kimi K2.61461
25Claude Sonnet 4.61460
26Qwen3.6 Max Preview1460
27Gemini 3.5 Flash Lite1459
28Qwen3.7 Plus1458
29GLM 51457
30Hy31453
31Claude Sonnet 4.51452
32Gemma 4 31B1451
33Claude Opus 4.11449
34Grok 4.31446
35MiniMax M31445
36Gemini 2.5 Pro1444
37Qwen3.6 Plus1443
38Qwen3.5 397B A17B1442
39GLM 4.71442
40Inkling1442
41Gemma 4 26B A4B 1438
42MiMo-V2.51434
43DeepSeek V4 Flash1433
44GLM 5V Turbo1433
45Gemini 3.1 Flash Lite Preview1432
46Inkling Small1431
47Claude Opus 4.51430
48Mistral Medium 3.51427
49DeepSeek V3.21425
50GLM 4.61425
51DeepSeek V3.2 Exp1423
52Claude Opus 41420
53DeepSeek V3.11418
54Qwen3.5-122B-A10B1417
55MiniMax M2.71416
56o31415
57Qwen3 VL 235B A22B Instruct1415
58DeepSeek V3.1 Terminus1415
59Hy3 preview1412
60GLM 4.51411
61Qwen3.5-27B1408
62Qwen3 Next 80B A3B Instruct1401
63Qwen3.5-Flash1397
64Gemini 2.5 Flash1395
65Qwen3 VL 235B A22B Thinking1395
66Qwen3.5-35B-A3B1395
67Step 3.5 Flash1394
68GPT-5 Mini1390
69MiniMax M2.51390
70Claude Sonnet 41387
71GPT-4.1 Mini1383
72o4-mini1380
73DeepSeek R1-05281380
74GLM 4.6V1377
75GLM 4.5 Air1373
76o3-mini1371
77DeepSeek R11369
78Qwen3 Next 80B A3B Thinking1369
79Trinity Large Thinking1369
80GLM 4.7 Flash1368
81MiniMax M11364
82o3 Mini High1364
83Claude 3.7 Sonnet1354
84GLM 4.5V1354
85Gemini 2.0 Flash1352
86gpt-oss-120b1352
87o11350
88Qwen3 8B1347
89Mercury 21347
90MiniMax M21346
91DeepSeek V3 (March 2025)1345
92Grok 31342
93GPT-5 Nano1337
94Nova 2 Lite1337
95o1 Preview1334
96Llama 4 Maverick1325
97GPT-4.1 Nano1322
98DeepSeek V31318
99GPT-4o-mini (2024-07-18)1318
100gpt-oss-20b1317
101Mistral Large 24071314
102Granite 4.1 8B1306
103Olmo 3 32B Think1305
104o1-mini1304
105GPT-4.11300
106GPT-4o1286
107Gemini 1.5 Pro1281
108Mistral Large 21280
109GPT-41275
110Claude 3.5 Sonnet1271
111Qwen 2.5 Coder 32B1271
112Grok 21262
113Qwen 2.5 72B1261
114Command R+1261
115Command A1261
116Claude 3 Haiku1261
117Phi-41256
118GPT-4 Turbo1255
119Mistral Large 21250
120Command R+ (08-2024)1250
121Llama 3.3 70B1243
122Claude Haiku 4.51240
123Claude 3 Opus1232
124Llama 3.1 405B1229
125GPT-4o mini1222
126Llama 3.1 8B Instruct1211
127Llama 3.1 70B1198
128Claude 3.5 Haiku1178
129Llama 3.2 3B Instruct1166
130Mixtral 8x22B1146
131Llama 3.2 1B Instruct1111

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-08T06:30:09.792Z. Some scores are approximate.

Frequently Asked Questions

AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.

Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.

A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.

Arena AI Benchmarks - Compare Top Models | LM Market Cap