Skip to content

AI Model Speed Comparison

Compare AI model speeds across 300 models. See latency, tokens per second, and streaming capabilities to find the fastest models for your real-time applications.

300
Models Benchmarked
300
Streaming
300
TPS Data
13
Free

Speed Rankings - All Models

#ModelScore
1LFM2.5-2.6B (free)Liquid AI40
2Qwen3 30B A3B Thinking 2507Alibaba64
3DeepSeek V3.2DeepSeek83
4GLM 4.5Zhipu AI75
5R1 0528DeepSeek79
6GLM 4.5VZhipu AI62
7Muse Glimmer 30Bmeta40
8DeepSeek V3.1DeepSeek72
9gpt-oss-20b (batch)OpenAI57
10Hy-MT2-1.8BTencent40
11Phi 4Microsoft60
12Mixtral 8x22B InstructMistral AI63
13Qwen3 8BAlibaba61
14R1 Distill Llama 70BDeepSeek41
15Mistral Large 3 2512 (batch)Mistral AI67
16Qwen3 Next 80B A3B InstructAlibaba67
17GPT-5 MiniOpenAI54
18GLM 4.5 AirZhipu AI71
19GPT-4o-mini (2024-07-18)OpenAI57
20GLM 4.7 FlashZhipu AI64
21GLM 4.6Zhipu AI75
22GPT-5 Mini (batch)OpenAI77
23gpt-oss-120bOpenAI62
24Llama 3.1 8B InstructMeta45
25MiniMax M2-herMiniMax68
26GLM 4.7Zhipu AI75
27GPT-5.1-Codex-MiniOpenAI88
28Mercury 2.5Inception61
29Granite 4.2 8BIBM54
30Solar Pro 4Upstage40
31Gemma 4 31B (free)Google81
32Qwen3.5-35B-A3BAlibaba76
33Qwen3 30B A3BAlibaba64
34R1DeepSeek74
35Gemma 4 26B A4B Google73
36DeepSeek V3.1 TerminusDeepSeek69
37Llama 3.3 70B InstructMeta67
38Claude 3 HaikuAnthropic51
39Hy-MT2-30B-A3BTencent40
40GLM 5.2 (free)Zhipu AI76

Understanding AI Model Speed

Latency (Time to First Token)

How quickly the model starts responding. Critical for chatbots and interactive applications. Sub-500ms latency feels instantaneous to users.

Tokens Per Second (TPS)

Generation speed after the first token. Higher TPS means faster completion of long responses. Premium models typically range from 30-100+ TPS.

Streaming for Perceived Speed

Streaming shows tokens as they generate, making the model feel much faster. Even a slow model (20 TPS) feels responsive when streaming compared to waiting for a complete response.

Speed vs Quality Trade-off

Smaller models are generally faster. Budget models often have the best TPS. Reasoning models are slower but more accurate - choose based on your latency requirements.

Frequently Asked Questions

We measure speed using three metrics: tokens per second (generation throughput), time to first token (initial latency), and overall speed index (a weighted composite). These measurements come from real API calls against upstream provider endpoints and are refreshed regularly.

Latency directly impacts user experience. For real-time applications like chatbots and code completion, low time-to-first-token is critical. For batch processing, tokens-per-second throughput matters more. The right speed metric depends on your use case.

Speed varies by provider infrastructure and model size. Smaller models generally respond faster but may sacrifice quality. Our speed comparison tool shows real-time measurements across all models, helping you find the optimal speed-quality tradeoff.

AI Model Speed Test - Compare Response Times | LM Market Cap