Skip to content

AI Model Capacity & Efficiency Comparison

Compare context windows, output capacity, and cost efficiency across 406+ models. Data sourced live from upstream provider APIs.

Note: This page shows capacity specs and pricing aggregated from upstream provider APIs. Real-world latency and tokens-per-second vary by load, prompt length, and provider infrastructure. For speed benchmarks, see Artificial Analysis or the model provider's own documentation.

Largest Context

2M

Grok 4.20 Multi-Agent

Largest Output

1.0M

MoonshotAI Kimi Latest

Cheapest (non-free)

$0.01/1M in

Ling-2.6-flash

Best Output Ratio

100%

GPT-3.5 Turbo (older v0613)

Coding & LLM Models

ModelProviderContext WindowMax OutputOutput RatioInput $/1MOutput $/1MEfficiency (derived)Capabilities
Grok 4.20 Multi-AgentxAI2M$1.25$2.5048
Grok 4.20xAI2M$1.25$2.5048
Llama 4 ScoutMeta1.3M16K1%$0.10$0.3089
GPT-5.6 Luna ProOpenAI1.1M128K12%$0.10$0.6088
GPT-5.6 Luna Pro (batch)OpenAI1.1M128K12%$0.10$0.6088
GPT-5.6 LunaOpenAI1.1M128K12%$0.10$0.6088
GPT-5.6 Luna (batch)OpenAI1.1M128K12%$0.10$0.6088
GPT-5.6 Terra ProOpenAI1.1M128K12%$1.00$6.0050
GPT-5.6 Terra Pro (batch)OpenAI1.1M128K12%$1.00$6.0050
GPT-5.6 TerraOpenAI1.1M128K12%$1.00$6.0050
GPT-5.6 Terra (batch)OpenAI1.1M128K12%$1.00$6.0050
GPT-5.6 Sol ProOpenAI1.1M128K12%$5.00$30.0028
GPT-5.6 Sol Pro (batch)OpenAI1.1M128K12%$2.50$15.0036
GPT-5.6 SolOpenAI1.1M128K12%$5.00$30.0028
GPT-5.6 Sol (batch)OpenAI1.1M128K12%$2.50$15.0036
OpenAI GPT Latest~openai1.1M128K12%$5.00$30.0028
GPT-5.5 ProOpenAI1.1M128K12%$30.00$180.0017
GPT-5.5 Pro (batch)OpenAI1.1M128K12%$15.00$90.0020
GPT-5.5OpenAI1.1M128K12%$5.00$30.0028
GPT-5.5 (batch)OpenAI1.1M128K12%$2.50$15.0036
MiMo-V2.5-ProXiaomi1.1M131K12%$0.43$0.8766
MiMo-V2.5Xiaomi1.1M131K12%$0.14$0.2884
GPT-5.4 ProOpenAI1.1M128K12%$30.00$180.0017
GPT-5.4 Pro (batch)OpenAI1.1M128K12%$15.00$90.0020
GPT-5.4OpenAI1.1M128K12%$2.50$15.0036
GPT-5.4 (batch)OpenAI1.1M128K12%$1.25$7.5046
LongCat 2.0Meituan1.0M262K25%$0.30$1.2073
Muse Spark 1.2meta1.0M$1.25$4.2546
DeepSeek V4 Flash Latest~deepseek1.0M384K37%$0.09$0.1889
DeepSeek V4 Flash 0731DeepSeek1.0M384K37%$0.09$0.1889
Laguna S 2.1poolside1.0M131K13%$0.09$0.1889
Gemini 3.6 FlashGoogle1.0M66K6%$1.50$7.5043
Gemini 3.6 Flash (batch)Google1.0M66K6%$0.75$3.7555
Gemini 3.5 Flash LiteGoogle1.0M66K6%$0.30$2.5073
Gemini 3.5 Flash Lite (batch)Google1.0M66K6%$0.15$1.2583
Inklingthinkingmachines1.0M262K25%$0.95$4.0551
Kimi K3Moonshot AI1.0M$3.00$15.0033
Muse Spark 1.1meta1.0M$1.25$4.2546
GLM 5.2Zhipu AI1.0M128K12%$0.21$0.6579
MiniMax M3MiniMax1.0M512K49%$0.30$1.2073
Gemini 3.5 FlashGoogle1.0M66K6%$1.50$9.0043
Gemini 3.5 Flash (batch)Google1.0M66K6%$0.75$4.5055
Gemini 3.1 Flash LiteGoogle1.0M66K6%$0.25$1.5076
Gemini 3.1 Flash Lite (batch)Google1.0M66K6%$0.13$0.7585
Google Gemini Pro Latest~google1.0M66K6%$2.00$12.0039
MoonshotAI Kimi Latest~moonshotai1.0M1.0M100%$2.50$14.0036
Google Gemini Flash Latest~google1.0M66K6%$1.50$7.5043
DeepSeek V4 ProDeepSeek1.0M384K37%$0.43$0.8766
DeepSeek V4 Flash 0423DeepSeek1.0M393K38%$0.14$0.2884
Lyria 3 Pro PreviewGoogle1.0M66K6%FreeFree100
Lyria 3 Clip PreviewGoogle1.0M66K6%FreeFree100
Gemini 3.1 Flash Lite PreviewGoogle1.0M66K6%$0.25$1.5076
Gemini 3.1 Pro Preview Custom ToolsGoogle1.0M66K6%$2.00$12.0039
Gemini 3.1 Pro PreviewGoogle1.0M66K6%$2.00$12.0039
Gemini 3.1 Pro Preview (batch)Google1.0M66K6%$1.00$6.0050
Gemini 3 Flash PreviewGoogle1.0M66K6%$0.50$3.0063
Gemini 3 Flash Preview (batch)Google1.0M66K6%$0.25$1.5076
Gemini 2.5 Flash LiteGoogle1.0M66K6%$0.10$0.4088
Gemini 2.5 Flash Lite (batch)Google1.0M66K6%$0.05$0.2093
Gemini 2.5 FlashGoogle1.0M66K6%$0.30$2.5073
Gemini 2.5 Flash (batch)Google1.0M66K6%$0.15$1.2583
Gemini 2.5 ProGoogle1.0M66K6%$1.25$10.0046
Gemini 2.5 Pro (batch)Google1.0M66K6%$0.63$5.0059
Gemini 2.5 Pro Preview 06-05Google1.0M66K6%$1.25$10.0046
Gemini 2.5 Pro Preview 05-06Google1.0M66K6%$1.25$10.0046
Llama Guard 4 12BMeta1.0M16K2%$0.18$0.1881
Llama 4 MaverickMeta1.0M16K2%$0.20$0.8079
GPT-4.1OpenAI1.0M33K3%$2.00$8.0039
GPT-4.1 (batch)OpenAI1.0M33K3%$1.00$4.0050
GPT-4.1 MiniOpenAI1.0M33K3%$0.40$1.6067
GPT-4.1 Mini (batch)OpenAI1.0M33K3%$0.20$0.8079
GPT-4.1 NanoOpenAI1.0M33K3%$0.10$0.4088
GPT-4.1 Nano (batch)OpenAI1.0M33K3%$0.05$0.2093
Palmyra X5Writer1.0M8K1%$0.60$6.0060
MiniMax-01MiniMax1.0M1.0M100%$0.20$1.1079
Qwen3.8 MaxAlibaba1M131K13%$2.00$6.0039
Qwen3.7 FlashAlibaba1M66K7%$0.03$0.1396
Claude Opus 5 (Fast)Anthropic1M128K13%$10.00$50.0022
Claude Opus 5Anthropic1M128K13%$5.00$25.0028
Claude Opus 5 (batch)Anthropic1M128K13%$2.50$12.5035
Claude Sonnet 5Anthropic1M128K13%$2.00$10.0039
Claude Sonnet 5 (batch)Anthropic1M128K13%$1.00$5.0050
Fugu Ultrasakana1M128K13%$5.00$30.0028
Claude Fable Latest~anthropic1M128K13%$10.00$50.0022
Claude Fable 5Anthropic1M128K13%$10.00$50.0022
Claude Fable 5 (batch)Anthropic1M128K13%$5.00$25.0028
Nemotron 3 Ultra (free)NVIDIA1M66K7%FreeFree100
Qwen3.7 PlusAlibaba1M131K13%$0.32$1.2871
Claude Opus 4.8 (Fast)Anthropic1M128K13%$10.00$50.0022
Claude Opus 4.8Anthropic1M128K13%$5.00$25.0028
Claude Opus 4.8 (batch)Anthropic1M128K13%$2.50$12.5035
Qwen3.7 MaxAlibaba1M131K13%$1.48$4.4243
Claude Opus 4.7 (Fast)Anthropic1M128K13%$30.00$150.0017
Grok 4.3xAI1M$1.25$2.5046
Anthropic Claude Sonnet Latest~anthropic1M128K13%$2.00$10.0039
Qwen3.5 Plus 2026-04-20Alibaba1M66K7%$0.30$1.8072
Qwen3.6 FlashAlibaba1M66K7%$0.19$1.1380
Claude Opus Latest~anthropic1M128K13%$5.00$25.0028
Claude Opus 4.7Anthropic1M128K13%$5.00$25.0028
Claude Opus 4.7 (batch)Anthropic1M128K13%$2.50$12.5035
Qwen3.6 PlusAlibaba1M66K7%$0.33$1.9571
Nemotron 3 SuperNVIDIA1M16K2%$0.08$0.4089
Qwen3.5-FlashAlibaba1M66K7%$0.07$0.2691
Claude Sonnet 4.6Anthropic1M128K13%$3.00$15.0033
Claude Sonnet 4.6 (batch)Anthropic1M128K13%$1.50$7.5043
Qwen3.5 Plus 2026-02-15Alibaba1M66K7%$0.26$1.5675
Claude Opus 4.6Anthropic1M128K13%$5.00$25.0028
Claude Opus 4.6 (batch)Anthropic1M128K13%$2.50$12.5035
Nova 2 LiteAmazon1M66K7%$0.30$2.5072
Nova Premier 1.0Amazon1M32K3%$2.50$12.5035
Claude Sonnet 4.5Anthropic1M64K6%$3.00$15.0033
Claude Sonnet 4.5 (batch)Anthropic1M64K6%$1.50$7.5043
Qwen3 Coder PlusAlibaba1M66K7%$0.65$3.2558
Qwen3 Coder FlashAlibaba1M66K7%$0.20$0.9779
Qwen Plus 0728Alibaba1M33K3%$0.26$0.7875
Qwen Plus 0728 (thinking)Alibaba1M33K3%$0.40$1.2067
MiniMax M1MiniMax1M40K4%$0.55$2.2061
Claude Sonnet 4Anthropic1M64K6%$3.00$15.0033
Qwen-PlusAlibaba1M33K3%$0.26$0.7875
Inkling Smallthinkingmachines524K262K50%$0.45$1.2062
Inkling (batch)thinkingmachines524K$0.50$2.0260
MiniMax M3 (batch)MiniMax524K$0.15$0.6079
Nemotron 3 UltraNVIDIA512K$0.60$3.6057
Nemotron 3 Ultra (batch)NVIDIA512K$0.30$1.8069
GLM 5.2 (batch)Zhipu AI512K$0.70$2.2054
Grok 4.5xAI500K$2.00$6.0037
Grok Latest~x-ai500K$2.00$6.0037
GPT Chat LatestOpenAI400K128K32%$5.00$30.0026
OpenAI GPT Mini Latest~openai400K128K32%$0.75$4.5051
GPT-5.4 NanoOpenAI400K128K32%$0.20$1.2574
GPT-5.4 Nano (batch)OpenAI400K128K32%$0.10$0.6382
GPT-5.4 MiniOpenAI400K128K32%$0.75$4.5051
GPT-5.4 Mini (batch)OpenAI400K128K32%$0.38$2.2564
GPT-5.3-CodexOpenAI400K128K32%$1.75$14.0038
GPT-5.2-CodexOpenAI400K128K32%$1.75$14.0038
GPT-5.2 ProOpenAI400K128K32%$21.00$168.0017
GPT-5.2 Pro (batch)OpenAI400K128K32%$10.50$84.0021
GPT-5.2OpenAI400K128K32%$1.75$14.0038
GPT-5.2 (batch)OpenAI400K128K32%$0.88$7.0049
GPT-5.1-Codex-MaxOpenAI400K128K32%$1.25$10.0043
GPT-5.1OpenAI400K128K32%$1.25$10.0043
GPT-5.1 (batch)OpenAI400K128K32%$0.63$5.0055
GPT-5.1-CodexOpenAI400K128K32%$1.25$10.0043
GPT-5.1-Codex-MiniOpenAI400K128K32%$0.25$2.0070
GPT-5 ProOpenAI400K128K32%$15.00$120.0019
GPT-5 Pro (batch)OpenAI400K128K32%$7.50$60.0023
GPT-5 Codex (batch)OpenAI400K128K32%$0.63$5.0055
GPT-5OpenAI400K128K32%$1.25$10.0043
GPT-5 (batch)OpenAI400K128K32%$0.63$5.0055
GPT-5 MiniOpenAI400K128K32%$0.25$2.0070
GPT-5 Mini (batch)OpenAI400K128K32%$0.13$1.0080
GPT-5 NanoOpenAI400K128K32%$0.05$0.4087
GPT-5 Nano (batch)OpenAI400K128K32%$0.02$0.2090
Nova Lite 1.0Amazon300K5K2%$0.06$0.2484
Nova Pro 1.0Amazon300K5K2%$0.80$3.2049
Ling 3.0 Tiny (free)inclusionai262K33K13%FreeFree90
Ling-3.0-flashinclusionai262K33K13%$0.02$0.0687
Laguna S 2.1 (free)poolside262K33K13%FreeFree90
Hy3Tencent262K128K49%$0.13$0.5376
Laguna XS 2.1poolside262K33K13%$0.06$0.1283
Laguna XS 2.1 (free)poolside262K33K13%FreeFree90
Kimi K2.7 CodeMoonshot AI262K262K100%$0.70$3.5051
Kimi K2.7 Code (batch)Moonshot AI262K$0.47$2.0058
Step 3.7 FlashStepFun262K256K98%$0.20$1.1571
Ring-2.6-1Tinclusionai262K66K25%$0.07$0.6381
Mistral Medium 3.5Mistral AI262K$1.50$7.5039
Qwen3.6 35B A3BAlibaba262K262K100%$0.14$1.0076
Qwen3.6 Max PreviewAlibaba262K66K25%$1.03$6.1645
Qwen3.6 27BAlibaba262K262K100%$0.60$3.6054
Ling-2.6-1Tinclusionai262K33K13%$0.07$0.6381
Hy3 previewTencent262K$0.06$0.2183
Ling-2.6-flashinclusionai262K33K13%$0.01$0.0389
Kimi K2.6Moonshot AI262K262K100%$0.58$2.4454
Gemma 4 26B A4B Google262K16K6%$0.07$0.3482
Gemma 4 26B A4B (free)Google262K33K13%FreeFree90
Gemma 4 31BGoogle262K262K100%$0.10$0.3479
Gemma 4 31B (free)Google262K33K13%FreeFree90
Trinity Large Thinkingarcee-ai262K262K100%$0.22$0.8570
KAT-Coder-Pro V2Kuaishou262K80K31%$0.30$1.2065
Mistral Small 4Mistral AI262K$0.15$0.6075
Nemotron 3 Super (free)NVIDIA262K262K100%FreeFree90
Seed-2.0-LiteByteDance262K131K50%$0.25$2.0068
Qwen3.5-9BAlibaba262K262K100%$0.10$0.1579
Seed-2.0-MiniByteDance262K131K50%$0.10$0.4079
Qwen3.5-35B-A3BAlibaba262K262K100%$0.14$1.0076
Qwen3.5-27BAlibaba262K66K25%$0.20$1.5672
Qwen3.5-122B-A10BAlibaba262K82K31%$0.29$2.4066
Qwen3.5 397B A17BAlibaba262K66K25%$0.39$2.3461
Qwen3 Max ThinkingAlibaba262K66K25%$0.78$3.9049
Qwen3 Coder NextAlibaba262K262K100%$0.12$0.8077
Step 3.5 FlashStepFun262K66K25%$0.10$0.3079
Kimi K2.5Moonshot AI262K262K100%$0.57$2.8555
Seed 1.6 FlashByteDance262K33K13%$0.07$0.3081
Seed 1.6ByteDance262K33K13%$0.25$2.0068
Nemotron 3 Nano 30B A3BNVIDIA262K262K100%$0.05$0.2084
Ministral 3 14B 2512Mistral AI262K$0.20$0.2071
Ministral 3 8B 2512Mistral AI262K$0.15$0.1575
Mistral Large 3 2512Mistral AI262K$0.50$1.5057
Kimi K2 ThinkingMoonshot AI262K100K38%$0.60$2.5054
Qwen3 VL 8B InstructAlibaba262K33K13%$0.12$0.4578
Qwen3 VL 30B A3B ThinkingAlibaba262K33K13%$0.20$2.4071
Qwen3 VL 30B A3B InstructAlibaba262K16K6%$0.15$0.6075
Qwen3 VL 235B A22B InstructAlibaba262K33K13%$0.21$1.9071
Qwen3 MaxAlibaba262K66K25%$0.78$3.9049
Qwen3 Next 80B A3B ThinkingAlibaba262K262K100%$0.15$1.2075
Qwen3 Next 80B A3B InstructAlibaba262K16K6%$0.09$1.1080
Kimi K2 0905Moonshot AI262K100K38%$0.60$2.5054
Qwen3 Coder 30B A3B InstructAlibaba262K33K13%$0.07$0.2782
Qwen3 30B A3B Instruct 2507Alibaba262K32K12%$0.05$0.1984
Qwen3 235B A22B Thinking 2507Alibaba262K$0.23$2.3069
Qwen3 Coder 480B A35BAlibaba262K66K25%$0.30$1.0065
Qwen3 235B A22B Instruct 2507Alibaba262K16K6%$0.09$0.5580
Gemma 3 27BGoogle262K131K50%$0.08$0.4581
Falcon-H1-Arabic 34B InstructTII262K8K3%FreeFree90
Falcon-H1-Arabic 7B InstructTII262K8K3%FreeFree90
KAT-Coder-Air V2.5Kuaishou256K80K31%$0.15$0.6075
KAT-Coder-Pro V2.5Kuaishou256K80K31%$0.74$2.9650
North Mini Code (free)Cohere256K64K25%FreeFree90
Grok Build 0.1xAI256K$1.00$2.0045
Nemotron 3 Nano Omni (free)NVIDIA256K66K26%FreeFree90
Nemotron 3 Nano 30B A3B (free)NVIDIA256KFreeFree90
Jamba Large 1.7AI21 Labs256K4K2%$2.00$8.0035
Codestral 2508Mistral AI256K$0.30$0.9065
Mistral Small 3.2 24BMistral AI256K16K6%$0.09$0.2580
Command ACohere256K8K3%$2.50$10.0032
GLM 5.1Zhipu AI205K131K64%$0.95$2.9945
MiniMax M2.7MiniMax205K131K64%$0.27$1.0866
MiniMax M2.5MiniMax205K197K96%$0.22$0.9069
GLM 5Zhipu AI205K131K64%$0.95$2.5545
MiniMax M2.1MiniMax205K131K64%$0.30$1.2064
GLM 4.7Zhipu AI205K131K64%$0.40$1.7559
MiniMax M2MiniMax205K131K64%$0.26$1.0266
GLM 4.6Zhipu AI205K131K64%$0.50$2.0056
GLM 5V TurboZhipu AI203K131K65%$1.20$4.0041
GLM 5 TurboZhipu AI203K131K65%$1.20$4.0041
GLM 4.7 FlashZhipu AI203K16K8%$0.06$0.4081
Anthropic Claude Haiku Latest~anthropic200K64K32%$1.00$5.0044
Claude Opus 4.5Anthropic200K64K32%$5.00$25.0025
Claude Opus 4.5 (batch)Anthropic200K64K32%$2.50$12.5031
Sonar Pro SearchPerplexity200K8K4%$3.00$15.0029
Claude Haiku 4.5Anthropic200K64K32%$1.00$5.0044
Claude Haiku 4.5 (batch)Anthropic200K64K32%$0.50$2.5056
Claude Opus 4.1Anthropic200K32K16%$15.00$75.0018
Claude Opus 4.1 (batch)Anthropic200K32K16%$7.50$37.5022
o3 ProOpenAI200K100K50%$20.00$80.0016
o3 Pro (batch)OpenAI200K100K50%$10.00$40.0020
Claude Opus 4Anthropic200K32K16%$15.00$75.0018
o4 Mini HighOpenAI200K100K50%$1.10$4.4043
o4 Mini High (batch)OpenAI200K100K50%$0.55$2.2054
o3OpenAI200K100K50%$2.00$8.0034
o3 (batch)OpenAI200K100K50%$1.00$4.0044
o4 MiniOpenAI200K100K50%$1.10$4.4043
o4 Mini (batch)OpenAI200K100K50%$0.55$2.2054
o1-proOpenAI200K100K50%$150.00$600.0011
o1-pro (batch)OpenAI200K100K50%$75.00$300.0012
Sonar ProPerplexity200K8K4%$3.00$15.0029
o3 Mini HighOpenAI200K100K50%$1.10$4.4043
o3 Mini High (batch)OpenAI200K100K50%$0.55$2.2054
o3 MiniOpenAI200K100K50%$1.10$4.4043
o3 Mini (batch)OpenAI200K100K50%$0.55$2.2054
o1OpenAI200K100K50%$15.00$60.0018
o1 (batch)OpenAI200K100K50%$7.50$30.0022
Claude 3 HaikuAnthropic200K4K2%$0.25$1.2567
Composer 2Cursor200K66K33%$0.50$2.5056
Composer 2 FastCursor200K66K33%$1.50$7.5038
DeepSeek V3.2DeepSeek164K164K100%$0.26$0.3865
DeepSeek V3.2 ExpDeepSeek164K66K40%$0.27$0.4164
DeepSeek V3.1 TerminusDeepSeek164K33K20%$0.27$1.0064
DeepSeek V3.1DeepSeek164K33K20%$0.25$0.9566
R1 0528DeepSeek164K33K20%$0.50$2.1555
DeepSeek V3 0324DeepSeek164K66K40%$0.27$1.1264
R1DeepSeek164K16K10%$0.70$2.5049
DeepSeek V3DeepSeek164K16K10%$0.26$1.0365
Aion-3.0-Miniaion-labs131K33K25%$0.70$1.4048
Aion-3.0aion-labs131K33K25%$3.00$6.0028
Granite 4.1 8BIBM131K131K100%$0.05$0.1079
Aion-2.0aion-labs131K33K25%$0.80$1.6046
Solar Pro 3Upstage131K131K100%$0.15$0.6071
GLM 4.6VZhipu AI131K33K25%$0.30$0.9062
Ministral 3 3B 2512Mistral AI131K$0.10$0.1075
gpt-oss-safeguard-20bOpenAI131K66K50%$0.07$0.3077
Qwen3 VL 32B InstructAlibaba131K33K25%$0.10$0.4274
Qwen3 VL 8B ThinkingAlibaba131K33K25%$0.18$2.1069
Qwen3 VL 235B A22B ThinkingAlibaba131K33K25%$0.40$4.0057
Mistral Medium 3.1Mistral AI131K$0.40$2.0057
gpt-oss-120bOpenAI131K131K100%$0.04$0.1781
gpt-oss-20bOpenAI131K131K100%$0.03$0.1382
gpt-oss-20b (free)OpenAI131K33K25%FreeFree85
GLM 4.5Zhipu AI131K98K75%$0.60$2.2051
GLM 4.5 AirZhipu AI131K98K75%$0.13$0.8572
Kimi K2 0711Moonshot AI131K100K77%$0.57$2.3051
Hunyuan A13B InstructTencent131K131K100%$0.14$0.5771
Mistral Medium 3Mistral AI131K$0.40$2.0057
Virtuoso Largearcee-ai131K64K49%$0.75$1.2047
Qwen3 30B A3BAlibaba131K16K13%$0.12$0.5073
Qwen3 8BAlibaba131K8K6%$0.12$0.4573
Qwen3 14BAlibaba131K8K6%$0.23$0.9166
Qwen3 32BAlibaba131K16K13%$0.08$0.2877
Qwen3 235B A22BAlibaba131K8K6%$0.45$1.8255
Gemma 3 4BGoogle131K16K13%$0.05$0.1079
Gemma 3 12BGoogle131K16K13%$0.05$0.1579
Llama 3.3 70B InstructMeta131K16K13%$0.10$0.3275
Mistral Large 2407Mistral AI131K$2.00$6.0033
Llama 3.2 3B InstructMeta131K131K100%$0.05$0.3379
Llama 3.1 70B InstructMeta131K16K13%$0.40$0.4057
Llama 3.1 8B InstructMeta131K131K100%$0.05$0.0879
Mistral NemoMistral AI131K16K13%$0.02$0.0383
Falcon-H1-Arabic 3B InstructTII131K8K6%FreeFree85
Granite 4.0 MicroIBM131K131K100%$0.02$0.1183
Nemotron 3.5 Content Safety (free)NVIDIA128K8K6%FreeFree85
Mercury 2Inception128K50K39%$0.25$0.7564
GPT-5.3 ChatOpenAI128K16K13%$1.75$14.0034
GPT AudioOpenAI128K16K13%$2.50$10.0030
GPT Audio MiniOpenAI128K16K13%$0.60$2.4051
GPT-5.2 ChatOpenAI128K16K13%$1.75$14.0034
Cogito v2.1 671Bdeepcogito128K$1.25$1.2539
Nemotron Nano 12B 2 VL (free)NVIDIA128K128K100%FreeFree85
Nemotron Nano 9B V2 (free)NVIDIA128KFreeFree85
UI-TARS 7B ByteDance128K2K2%$0.10$0.2075
Mistral Small 3.1 24BMistral AI128K128K100%$0.35$0.5559
Sonar Reasoning ProPerplexity128K$2.00$8.0033
Sonar Deep ResearchPerplexity128K$2.00$8.0033
Qwen2.5 VL 72B InstructAlibaba128K$0.25$0.7564
Command R7B (12-2024)Cohere128K4K3%$0.04$0.1581
Nova Micro 1.0Amazon128K5K4%$0.04$0.1481
GPT-4o (2024-11-20)OpenAI128K16K13%$2.50$10.0030
Command R (08-2024)Cohere128K4K3%$0.15$0.6071
Command R+ (08-2024)Cohere128K4K3%$2.50$10.0030
GPT-4o (2024-08-06)OpenAI128K16K13%$2.50$10.0030
GPT-4o-miniOpenAI128K16K13%$0.15$0.6071
GPT-4o-mini (2024-07-18)OpenAI128K16K13%$0.15$0.6071
GPT-4o-mini (batch)OpenAI128K16K13%$0.07$0.3077
GPT-4oOpenAI128K16K13%$2.50$10.0030
GPT-4o (2024-05-13)OpenAI128K4K3%$5.00$15.0024
GPT-4o (batch)OpenAI128K16K13%$1.25$5.0039
GPT-4 TurboOpenAI128K4K3%$10.00$30.0019
GPT-4 Turbo (batch)OpenAI128K4K3%$5.00$15.0024
Mistral LargeMistral AI128K$2.00$6.0033
GPT-4 Turbo PreviewOpenAI128K4K3%$10.00$30.0019
SonarPerplexity127K$1.00$1.0042
ERNIE 4.5 VL 424B A47B Baidu123K16K13%$0.42$1.2556
Qwen3 30B A3B Thinking 2507Alibaba82K33K40%$0.20$2.4065
MiniMax M2-herMiniMax66K2K3%$0.30$1.2058
Olmo 3 32B ThinkAllen AI66K66K100%$0.15$0.5067
GLM 4.5VZhipu AI66K16K25%$0.60$1.8048
Reka Flash 3rekaai66K66K100%$0.10$0.2070
Mixtral 8x22B InstructMistral AI66K$2.00$6.0031
WizardLM-2 8x22BMicrosoft66K8K12%$0.62$0.6247
Llama 3.2 1B InstructMeta60K60K100%$0.03$0.2076
Perceptron Mk1perceptron33K8K25%$0.15$1.5062
Gemma 3n 4BGoogle33K$0.06$0.1269
SabaMistral AI33K$0.20$0.6059
Mistral Small 3Mistral AI33K16K50%$0.05$0.0870
Qwen2.5 Coder 32B InstructAlibaba33K33K100%$0.66$1.0043
Qwen2.5 7B InstructAlibaba33K33K100%$0.10$0.2066
Qwen2.5 72B InstructAlibaba33K16K50%$0.36$0.4052
Falcon Arabic 7B InstructTII33K8K25%FreeFree75
Falcon3 10B InstructTII33K8K25%FreeFree75
Falcon3 7B InstructTII33K8K25%FreeFree75
Falcon Mamba 7B InstructTII33K8K25%FreeFree75
Voxtral Small 24B 2507Mistral AI32K$0.10$0.3066
GPT-3.5 Turbo 16kOpenAI16K4K25%$3.00$4.0023
GPT-3.5 TurboOpenAI16K4K25%$0.50$1.5044
GPT-3.5 Turbo (batch)OpenAI16K4K25%$0.25$0.7553
Reka Edgerekaai16K16K100%$0.10$0.1062
Phi 4Microsoft16K16K100%$0.07$0.1464
R1 Distill Llama 70BDeepSeek8K8K100%$0.80$0.8035
Gemma 2 27BGoogle8K2K25%$0.65$0.6538
GPT-4OpenAI8K4K50%$30.00$60.0011
ALLaM 7B Instruct (preview)HUMAIN4K4K100%FreeFree60
ALLaM 1 13B InstructHUMAIN4K4K100%$1.80$1.8024
ALLaM 2 7B InstructHUMAIN4K4K100%FreeFree60
ALLaM 34BHUMAIN4K4K100%FreeFree60
GPT-3.5 Turbo (older v0613)OpenAI4K4K100%$1.00$2.0030
GPT-3.5 Turbo InstructOpenAI4K4K100%$1.50$2.0026
SWE-1.5WindsurfFreeFree0
autofixer-01VercelFreeFree0
MellumJetBrainsFreeFree0
Capabilities:VisionFunctionsStreamingJSON ModeReasoningWeb SearchImage Output

Image Generation Models

ModelProviderInput $/1MOutput $/1MCapabilities
GPT-5 ImageOpenAI$10.00$10.00
GPT-5.4 Image 2OpenAI$8.00$15.00
GPT-5 Image MiniOpenAI$2.50$2.00
Nano Banana Pro (Gemini 3 Pro Image)Google$2.00$12.00
Nano Banana Pro (Gemini 3 Pro Image Preview)Google$2.00$12.00
Nano Banana 2 (Gemini 3.1 Flash Image)Google$0.50$3.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google$0.50$3.00
Nano Banana (Gemini 2.5 Flash Image)Google$0.30$2.50
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google$0.25$1.50
Midjourney v6.1MidjourneyFreeFree
DALL-E 3OpenAIFree$40000.00
Stable Diffusion 3.5Stability AIFree$35000.00
FLUX.1 ProBlack Forest LabsFree$50000.00
Ideogram 2.0IdeogramFree$80000.00
Recraft V3RecraftFree$40000.00
Imagen 3GoogleFree$40000.00
Adobe Firefly 3AdobeFreeFree
Leonardo PhoenixLeonardo AIFreeFree
Frequently Asked Questions

AI speed is measured by time-to-first-token (TTFT) and tokens-per-second (TPS). TTFT measures how quickly the model starts responding. TPS measures how fast it generates output. Both matter for different use cases.

Speed varies by provider. Groq-hosted Llama achieves the fastest inference. Among major providers, Gemini Flash and GPT-4o Mini are consistently fast. Reasoning models like o3 and R1 are intentionally slower for better accuracy.

Smaller, faster models may sacrifice some quality. However, provider optimizations (quantization, speculative decoding) can speed up models without quality loss. The same model runs at different speeds on different providers.

AI Model Capacity & Efficiency Comparison | LM Market Cap