Skip to content

AI Model Speed Comparison

Compare AI model speeds across 300 models. See latency, tokens per second, and streaming capabilities to find the fastest models for your real-time applications.

300
Models Benchmarked
300
Streaming
300
TPS Data
17
Free

Speed Rankings - All Models

#ModelScore
1Reka Edgerekaai40
2Perceptron Mk1perceptron40
3Qwen3 30B A3B Thinking 2507Alibaba64
4DeepSeek V3.2DeepSeek81
5GLM 4.5Zhipu AI75
6R1 0528DeepSeek79
7GLM 4.5VZhipu AI63
8DeepSeek V3.1DeepSeek72
9Solar Pro 3Upstage40
10Falcon-H1-Arabic 3B InstructTII40
11Phi 4Microsoft60
12Mixtral 8x22B InstructMistral AI63
13Qwen3 8BAlibaba61
14Olmo 3 32B ThinkAllen AI55
15R1 Distill Llama 70BDeepSeek41
16Nemotron 3 Nano 30B A3B (free)NVIDIA40
17Kimi K2 ThinkingMoonshot AI53
18Qwen3 Next 80B A3B InstructAlibaba67
19GPT-5 MiniOpenAI64
20GLM 4.5 AirZhipu AI71
21GPT-4o-mini (2024-07-18)OpenAI57
22Falcon-H1-Arabic 7B InstructTII40
23Inkling Smallthinkingmachines72
24GLM 4.7 FlashZhipu AI64
25Mistral Large 3 2512Mistral AI67
26GLM 4.6Zhipu AI75
27GPT-5 Mini (batch)OpenAI77
28gpt-oss-120bOpenAI40
29gpt-oss-20b (free)OpenAI57
30Llama 3.1 8B InstructMeta45
31MiniMax M2-herMiniMax69
32GLM 4.7Zhipu AI75
33GPT-5.1-Codex-MiniOpenAI88
34Gemma 4 31B (free)Google81
35Qwen3.5-35B-A3BAlibaba76
36Aion-2.0aion-labs40
37Nemotron 3 Nano 30B A3BNVIDIA40
38Qwen3 30B A3BAlibaba64
39o3 Mini High (batch)OpenAI74
40Nemotron 3.5 Content Safety (free)NVIDIA40

Understanding AI Model Speed

Latency (Time to First Token)

How quickly the model starts responding. Critical for chatbots and interactive applications. Sub-500ms latency feels instantaneous to users.

Tokens Per Second (TPS)

Generation speed after the first token. Higher TPS means faster completion of long responses. Premium models typically range from 30-100+ TPS.

Streaming for Perceived Speed

Streaming shows tokens as they generate, making the model feel much faster. Even a slow model (20 TPS) feels responsive when streaming compared to waiting for a complete response.

Speed vs Quality Trade-off

Smaller models are generally faster. Budget models often have the best TPS. Reasoning models are slower but more accurate - choose based on your latency requirements.

Frequently Asked Questions

We measure speed using three metrics: tokens per second (generation throughput), time to first token (initial latency), and overall speed index (a weighted composite). These measurements come from real API calls against upstream provider endpoints and are refreshed regularly.

Latency directly impacts user experience. For real-time applications like chatbots and code completion, low time-to-first-token is critical. For batch processing, tokens-per-second throughput matters more. The right speed metric depends on your use case.

Speed varies by provider infrastructure and model size. Smaller models generally respond faster but may sacrifice quality. Our speed comparison tool shows real-time measurements across all models, helping you find the optimal speed-quality tradeoff.

AI Model Speed Test - Compare Response Times | LM Market Cap