Skip to content

Best AI Models for Multilingual Tasks

AI models ranked by multilingual performance using MMLU benchmark scores across languages. Find the best LLM for translation and non-English tasks.

Last updated: 11m ago
#1 Model

GPT-5.4

Score: 95.4

Average Score

74.5

Across all ranked models

Models Ranked

113

With benchmark data

Weights:MMLU (60%)Arena Elo (40%)

Top Best for Multilingual Models by Weighted Score

Top 15 models by weighted score

LMMarketCap.com

Benchmark Breakdown

Per-benchmark scores for top 10 models

MMLU
Arena Elo
LMMarketCap.com
#ModelScore
1GPT-5.4OpenAI95.4
2Claude Opus 4.6Anthropic95.3
3Gemini 3.1 Pro PreviewGoogle95.2
4GPT-5.2OpenAI94.8
5GPT-5.1OpenAI94.3
6GPT-5.5OpenAI93.8
7GPT-5OpenAI93.5
8Claude Sonnet 4.6Anthropic92.1
9Claude Sonnet 4.5Anthropic91.3
10Gemini 3 Flash PreviewGoogle91.1
11Gemini 2.5 ProGoogle90.7
12Claude Opus 4.5Anthropic90.2
13o3OpenAI89.7
14Claude Opus 4Anthropic89.3
15DeepSeek V3.2DeepSeek88.1
16R1 0528DeepSeek86.9
17Claude Sonnet 4Anthropic86.2
18R1DeepSeek85.7
19o1OpenAI85.1
20Claude Fable 5Anthropic85
21Qwen3.8 MaxAlibaba84.6
22Gemini 2.5 FlashGoogle84.5
23Claude Opus 4.7Anthropic83.7
24o3 MiniOpenAI83.5
25Muse Spark 1.1meta83.2
26DeepSeek V3 0324DeepSeek83.2
27Gemini 3.6 FlashGoogle82.9
28Claude Opus 4.8Anthropic82
29Gemini 3.5 FlashGoogle81.7
30GPT-5.2 ChatOpenAI81.6
31Llama 4 MaverickMeta81.1
32DeepSeek V3DeepSeek81
33Grok 4.5xAI80.5
34GLM 5.1Zhipu AI80.5
35MiMo-V2.5-ProXiaomi80.3
36GPT-4.1OpenAI80.2
37DeepSeek V4 ProDeepSeek79.8
38Kimi K2.6Moonshot AI79.5
39Qwen3.6 Max PreviewAlibaba79.3
40Gemini 3.5 Flash LiteGoogle79.2
41Qwen3.7 PlusAlibaba79.1
42GPT-4oOpenAI79
43GLM 5Zhipu AI78.9
44GPT-5.5 ProOpenAI78.5
45Hy3Tencent78.3
46Gemma 4 31BGoogle78.1
47Claude Opus 4.1Anthropic77.8
48MiniMax M3MiniMax77.2
49Qwen3.6 PlusAlibaba76.9
50Qwen3.5 397B A17BAlibaba76.8
51GLM 4.7Zhipu AI76.8
52Inklingthinkingmachines76.8
53Gemma 4 26B A4B Google76.2
54Mistral LargeMistral AI76.2
55MiMo-V2.5Xiaomi75.6
56GPT-4 TurboOpenAI75.6
57GLM 5V TurboZhipu AI75.5
58Gemini 3.1 Flash Lite PreviewGoogle75.4
59Inkling Smallthinkingmachines75.2
60Mistral Medium 3.5Mistral AI74.7
61Phi 4Microsoft74.6
62Llama 3.3 70B InstructMeta74.6
63GLM 4.6Zhipu AI74.4
64DeepSeek V3.2 ExpDeepSeek74.1
65DeepSeek V3.1DeepSeek73.4
66Claude Haiku 4.5Anthropic73.4
67Qwen3.5-122B-A10BAlibaba73.2
68MiniMax M2.7MiniMax73.1
69Qwen3 VL 235B A22B InstructAlibaba73
70DeepSeek V3.1 TerminusDeepSeek73
71Hy3 previewTencent72.5
72GLM 4.5Zhipu AI72.4
73Qwen3.5-27BAlibaba72
74Llama 3.1 70B InstructMeta71.5
75Qwen3 Next 80B A3B InstructAlibaba71
76GPT-4o-miniOpenAI70.7
77Qwen3.5-FlashAlibaba70.4
78Qwen3 VL 235B A22B ThinkingAlibaba70.1
79Qwen3.5-35B-A3BAlibaba70.1
80Step 3.5 FlashStepFun70
81GPT-5 MiniOpenAI69.4
82MiniMax M2.5MiniMax69.4
83GPT-4.1 MiniOpenAI68.4
84o4 MiniOpenAI68
85Llama 4 ScoutMeta67.7
86GLM 4.6VZhipu AI67.6
87GLM 4.5 AirZhipu AI67
88Qwen3 Next 80B A3B ThinkingAlibaba66.4
89Trinity Large Thinkingarcee-ai66.4
90GLM 4.7 FlashZhipu AI66.3
91MiniMax M1MiniMax65.7
92o3 Mini HighOpenAI65.7
93GLM 4.5VZhipu AI64.3
94gpt-oss-120bOpenAI64
95Gemma 2 27BGoogle63.9
96Qwen3 8BAlibaba63.3
97Mercury 2Inception63.3
98MiniMax M2MiniMax63.2
99GPT-5 NanoOpenAI61.9
100Nova 2 LiteAmazon61.9
101GPT-4.1 NanoOpenAI59.8
102GPT-4o-mini (2024-07-18)OpenAI59.2
103gpt-oss-20bOpenAI59.1
104Mistral Large 2407Mistral AI58.7
105Granite 4.1 8BIBM57.5
106Olmo 3 32B ThinkAllen AI57.4
107GPT-4OpenAI53.1
108Command ACohere51.1
109Claude 3 HaikuAnthropic51.1
110Command R+ (08-2024)Cohere49.6
111Llama 3.1 8B InstructMeta44.1
112Llama 3.2 3B InstructMeta37.7
113Llama 3.2 1B InstructMeta29.9

How scores are calculated

Each model's score is a weighted average of its available benchmark results. When a model is missing some benchmarks, the weights are re-normalized across the benchmarks that are available. All scores are on a 0-100 scale. Data sourced from official model cards, published papers, and third-party evaluation platforms.

Frequently Asked Questions

Based on our benchmark analysis, GPT-5.4 by OpenAI is currently the #1 ranked model for multilingual, with a weighted score of 95.4/100.

Models are ranked using a weighted average of MMLU, Arena Elo benchmark scores. All scores are normalized to a 0-100 scale.

We currently rank 113 models that have relevant benchmark data for multilingual tasks.

Best AI Models for Multilingual Tasks (2026) | LM Market Cap