Math 基准测试
Compare top models across the benchmark suite that best represents math performance. Use this page as the fastest way to inspect the relevant tests, then jump into the full matrix when you want broader context.
3
类别中的基准测试
52
有覆盖的模型
0
有人类基准的基准测试
1
饱和的基准测试
包含内容
The current benchmark set in this category, with context on what each test captures.
MATH Benchmark (500-problem subset)
Tests genuine mathematical reasoning, not just pattern matching. Reasoning models (o1, R1) dramatically outperform standard models here.
American Invitational Mathematics Examination 2024
Tests mathematical reasoning at competition level. Reasoning models achieve 70-90% while standard models struggle below 30%. Best differentiator for math ability.
Grade School Math 8K
饱和Useful baseline for math ability, but now saturated — top models exceed 95%. More useful for evaluating smaller or open-source models.
MATH Benchmark (500-problem subset)
Competition-level mathematics across algebra, geometry, number theory, counting/probability, intermediate algebra, and precalculus. 500-problem subset used by OpenAI for consistent evaluation.
Why it matters
Tests genuine mathematical reasoning, not just pattern matching. Reasoning models (o1, R1) dramatically outperform standard models here.
MATH-500 Scores (49 models)
| # | Model | Score |
|---|---|---|
| 1 | 🥇o3 | 99.0% |
| 2 | 🥈o3-mini | 97.9% |
| 3 | 🥉DeepSeek R1-0528 | 97.8% |
| 4 | DeepSeek R1 | 97.3% |
| 5 | o4-mini | 97.3% |
| 6 | o1 | 96.4% |
| 7 | Gemini 3 Pro | 96.0% |
| 8 | GPT-5.4 | 95.5% |
| 9 | Gemini 2.5 Pro | 95.2% |
| 10 | Grok 4 | 95.0% |
| 11 | o1 Preview | 94.8% |
| 12 | GPT-5.2 | 94.0% |
| 13 | GPT-5.1 | 93.5% |
| 14 | GPT-5 | 92.5% |
| 15 | DeepSeek V3 (March 2025) | 92.0% |
| 16 | Claude Opus 4.6 | 90.5% |
| 17 | DeepSeek V3 | 90.2% |
| 18 | o1-mini | 90.0% |
| 19 | Gemini 2.0 Flash | 89.7% |
| 20 | Claude Opus 4.5 | 88.1% |
| 21 | Gemini 3 Flash | 88.0% |
| 22 | Claude Opus 4 | 86.0% |
| 23 | Gemini 2.5 Flash | 85.8% |
| 24 | Claude Sonnet 4.6 | 85.3% |
| 25 | Grok 3 | 85.0% |
| 26 | Qwen 2.5 Coder 32B | 83.5% |
| 27 | Qwen 2.5 72B | 83.1% |
| 28 | Claude Sonnet 4.5 | 83.0% |
| 29 | Claude 3.7 Sonnet | 82.2% |
| 30 | Claude Sonnet 4 | 81.4% |
| 31 | Llama 4 Maverick | 81.0% |
| 32 | Phi-4 | 80.4% |
| 33 | GPT-4.1 | 78.5% |
| 34 | Claude 3.5 Sonnet | 78.3% |
| 35 | Llama 3.3 70B | 77.0% |
| 36 | GPT-4o | 76.6% |
| 37 | Grok 2 | 76.1% |
| 38 | Mistral Large 2 | 76.0% |
| 39 | Mistral Large 2 | 76.0% |
| 40 | Llama 3.1 405B | 73.8% |
| 41 | GPT-4 Turbo | 72.6% |
| 42 | Claude Haiku 4.5 | 72.5% |
| 43 | GPT-4o mini | 70.2% |
| 44 | Claude 3.5 Haiku | 69.2% |
| 45 | Llama 3.1 70B | 68.0% |
| 46 | Gemini 1.5 Pro | 67.7% |
| 47 | Claude 3 Opus | 60.1% |
| 48 | Mixtral 8x22B | 60.0% |
| 49 | Llama 4 Scout | 50.3% |
How to Read This Page
Performance Tiers
Model Types
Saturated benchmarks have top models clustered above 90%, making them less useful for comparison.
Scores sourced from official model cards, technical reports, and third-party evaluations (Artificial Analysis, LMSYS Arena). Last updated: 2026-08-08T06:30:09.792Z. Some scores are approximate.
AI benchmarks are grouped into categories like coding, math, reasoning, knowledge, and safety. Each category contains multiple standardized tests that measure specific aspects of model performance. This page focuses on one category so you can compare models within a specific skill area.
Each benchmark has its own scoring method - accuracy percentage, pass rate, Elo rating, or normalized score. We display raw scores from official evaluations and community-run tests. Scores are updated hourly as new evaluation results become available.
A saturated benchmark is one where top models score near the maximum (typically above 95%). This means the benchmark no longer effectively differentiates between the best models, and newer, harder benchmarks are needed to measure progress.