Best AI Models for Data Analysis
AI models ranked by data analysis capability using MMLU, MATH-500, and GPQA benchmark scores. Find the top LLM for data science, analytics, and insights.
o3
Score: 85.7
69.9
Across all ranked models
40
With benchmark data
Top Best for Data Analysis Models by Weighted Score
Top 15 models by weighted score
Benchmark Breakdown
Per-benchmark scores for top 10 models
| # | Model | Score |
|---|---|---|
| 1 | o3OpenAI | 85.7 |
| 2 | GPT-5.4OpenAI | 85.2 |
| 3 | R1 0528DeepSeek | 84.8 |
| 4 | o1OpenAI | 84.4 |
| 5 | GPT-5.2OpenAI | 84.3 |
| 6 | R1DeepSeek | 84.2 |
| 7 | GPT-5.1OpenAI | 84 |
| 8 | GPT-5OpenAI | 83.5 |
| 9 | Gemini 2.5 ProGoogle | 83.4 |
| 10 | o3 MiniOpenAI | 82.5 |
| 11 | Claude Opus 4.6Anthropic | 82.3 |
| 12 | DeepSeek V3 0324DeepSeek | 81.4 |
| 13 | Claude Opus 4.5Anthropic | 81 |
| 14 | DeepSeek V3DeepSeek | 80.3 |
| 15 | Claude Sonnet 4.6Anthropic | 79.8 |
| 16 | Gemini 3 Flash PreviewGoogle | 79.2 |
| 17 | Claude Sonnet 4.5Anthropic | 78.7 |
| 18 | o4 MiniOpenAI | 77.8 |
| 19 | Claude Sonnet 4Anthropic | 77.4 |
| 20 | Gemini 2.5 FlashGoogle | 77.2 |
| 21 | Llama 4 MaverickMeta | 76.5 |
| 22 | GPT-4.1OpenAI | 76.2 |
| 23 | GPT-4oOpenAI | 75.2 |
| 24 | Gemini 3.1 Pro PreviewGoogle | 74.1 |
| 25 | Llama 3.3 70B InstructMeta | 74.1 |
| 26 | GPT-5.5OpenAI | 73.9 |
| 27 | GPT-5.5 ProOpenAI | 73.9 |
| 28 | Mistral LargeMistral AI | 72.9 |
| 29 | GPT-4 TurboOpenAI | 72.5 |
| 30 | Claude Haiku 4.5Anthropic | 71.4 |
| 31 | DeepSeek V3.2DeepSeek | 70.8 |
| 32 | Llama 3.1 70B InstructMeta | 70.5 |
| 33 | GPT-4o-miniOpenAI | 69.2 |
| 34 | Phi 4Microsoft | 64.3 |
| 35 | Llama 4 ScoutMeta | 60.3 |
| 36 | Gemma 2 27BGoogle | 60.2 |
| 37 | Command R7B (12-2024)Cohere | 6.3 |
| 38 | Llama 3.1 8B InstructMeta | 5.9 |
| 39 | Llama 3.2 3B InstructMeta | 4.9 |
| 40 | Qwen2.5 7B InstructAlibaba | 4.4 |
How scores are calculated
Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.
Other Specialty Leaderboards
Based on our benchmark analysis, o3 by OpenAI is currently the #1 ranked model for data analysis, with a weighted score of 85.7/100.
Models are ranked using a weighted average of MMLU, MATH-500, GPQA benchmark scores. All scores are normalized to a 0-100 scale.
We currently rank 40 models that have relevant benchmark data for data analysis tasks.