Skip to content
基准测试类别

Arena 基准测试

Compare top models across the benchmark suite that best represents arena performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.

2

类别中的基准测试

135

有覆盖的模型

0

有人类基准的基准测试

0

饱和的基准测试

包含内容

The current benchmark set in this category, with context on what each test captures.

所有基准测试

LMSYS Chatbot Arena Elo Rating

The most trusted 'vibes-based' benchmark — reflects real human preferences, not just academic metrics. Widely considered the most meaningful overall ranking.

Elo rating

LiveBench (Dynamic)

Contamination-free by design — uses new questions regularly. Top models still score below 70%, making it highly discriminating.

average score %
Model type:

LMSYS Chatbot Arena Elo Rating

Human preference rating from 6M+ crowdsourced blind head-to-head comparisons. Users chat with two anonymous models and pick the better response.

Metric: Elo rating

Why it matters

The most trusted 'vibes-based' benchmark — reflects real human preferences, not just academic metrics. Widely considered the most meaningful overall ranking.

Arena Elo Scores (129 models)

Standard Reasoning Hybrid
LMMarketCap.com
#ModelScore
1🥇Claude Opus 4.61503
2🥈Gemini 3.1 Pro Preview1494
3🥉Gemini 3.1 Pro1494
4Muse Spark 1.11493
5Claude Opus 4.71491
6Gemini 3 Pro1486
7GPT-5.41485
8GPT-5.21481
9Claude Opus 4.81479
10Gemini 3.5 Flash1477
11GPT-5.2 Chat1476
12GPT-5.11475
13GPT-5.51475
14GLM 5.3 Flash1475
15Gemini 3 Flash1474
16Grok 4.51468
17MiMo-V2.5-Pro1467
18GLM 5.11466
19GPT-51465
20DeepSeek V4 Pro1463
21Grok 41462
22Claude Sonnet 4.61460
23Kimi K2.61460
24Qwen3.6 Max Preview1460
25GLM 51458
26Gemini 3.5 Flash Lite1456
27Hy31456
28Qwen3.7 Plus1456
29Claude Sonnet 4.51452
30Gemma 4 31B1451
31Claude Opus 4.11450
32Grok 4.31446
33Gemini 2.5 Pro1444
34Qwen3.6 Plus1443
35Qwen3.5 397B A17B1442
36GLM 4.71442
37MiniMax M31441
38Inkling1440
39Gemma 4 26B A4B 1438
40Qwen3.8 27B1437
41MiMo-V2.51434
42DeepSeek V4 Flash1433
43GLM 5V Turbo1433
44Gemini 3.1 Flash Lite Preview1432
45Claude Opus 4.51430
46Mistral Medium 3.51426
47DeepSeek V3.21425
48GLM 4.61425
49DeepSeek V3.2 Exp1422
50Claude Opus 41420
51DeepSeek V3.11417
52Qwen3.5-122B-A10B1417
53o31415
54MiniMax M2.71415
55DeepSeek V3.1 Terminus1415
56Qwen3 VL 235B A22B Instruct1414
57Hy3 preview1413
58GLM 4.51411
59Qwen3.5-27B1408
60Inkling Small1405
61Qwen3 Next 80B A3B Instruct1399
62Qwen3.5-Flash1397
63Gemini 2.5 Flash1395
64Qwen3 VL 235B A22B Thinking1395
65Qwen3.5-35B-A3B1395
66Step 3.5 Flash1394
67MiniMax M2.51390
68GPT-5 Mini1390
69Claude Sonnet 41387
70GPT-4.1 Mini1383
71o4-mini1380
72DeepSeek R1-05281380
73GLM 4.6V1378
74GLM 4.5 Air1373
75o3-mini1371
76DeepSeek R11369
77Qwen3 Next 80B A3B Thinking1369
78Trinity Large Thinking1369
79GLM 4.7 Flash1366
80MiniMax M11364
81o3 Mini High1363
82Claude 3.7 Sonnet1354
83GLM 4.5V1353
84Gemini 2.0 Flash1352
85gpt-oss-120b1352
86o11350
87Qwen3 8B1347
88Mercury 21346
89MiniMax M21346
90DeepSeek V3 (March 2025)1345
91Grok 31342
92GPT-5 Nano1337
93Nova 2 Lite1336
94o1 Preview1334
95Llama 4 Maverick1325
96GPT-4.1 Nano1322
97DeepSeek V31318
98GPT-4o-mini (2024-07-18)1318
99gpt-oss-20b1317
100Mistral Large 24071314
101o1-mini1304
102GPT-4.11300
103Granite 4.2 8B1287
104GPT-4o1286
105Gemini 1.5 Pro1281
106Mistral Large 21280
107GPT-41276
108Claude 3.5 Sonnet1271
109Qwen 2.5 Coder 32B1270
110Grok 21262
111Command R+1262
112Qwen 2.5 72B1261
113Command A1261
114Claude 3 Haiku1261
115Phi-41256
116GPT-4 Turbo1255
117Mistral Large 21250
118Command R+ (08-2024)1250
119Llama 3.3 70B1243
120Claude Haiku 4.51240
121Claude 3 Opus1232
122Llama 3.1 405B1229
123GPT-4o mini1222
124Llama 3.1 8B Instruct1211
125Llama 3.1 70B1198
126Claude 3.5 Haiku1178
127Llama 3.2 3B Instruct1167
128Mixtral 8x22B1146
129Llama 3.2 1B Instruct1111

How to Read This Page

Performance Tiers

Elite - Top 10% of the score range
Strong - Top 25% of the score range
Good - Above the midpoint
Below Average - Below the midpoint

Model Types

Standard - Direct inference, no chain-of-thought
Reasoning - Extended thinking (o1, R1) - slower but excels on math/reasoning
Hybrid - Optional thinking mode (Claude 3.7, Gemini 2.5) - can switch between fast and deep

Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.

Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-09-22T18:30:19.136Z. Some scores are approximate.

Frequently Asked Questions

AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.

Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.

A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.

Arena AI Benchmarks - Compare Top Models | LM Market Cap