AI Model Pricing History
Track how AI model API pricing changes over time. Covering 469+ models across 68 providers, with weekly snapshots since 2026-02-16. For the dispersion view of how tightly frontier prices cluster, see the LMC Price Convergence Index.
Average Pricing by Provider (Current)
Average input and output cost per 1M tokens across each provider's paid models.
| Provider | Models | Avg Input $/M | Avg Output $/M |
|---|---|---|---|
| inclusionai | 3 | $0.034 | $0.102 |
| poolside | 2 | $0.075 | $0.150 |
| rekaai | 2 | $0.100 | $0.150 |
| IBM | 2 | $0.038 | $0.181 |
| inference-net | 2 | $0.040 | $0.190 |
| Meta | 8 | $0.137 | $0.308 |
| Microsoft | 2 | $0.345 | $0.380 |
| Upstage | 3 | $0.097 | $0.387 |
| Inception | 2 | $0.145 | $0.450 |
| prism-ml | 1 | $0.075 | $0.500 |
| Tencent | 7 | $0.204 | $0.681 |
| NVIDIA | 5 | $0.200 | $0.690 |
| StepFun | 2 | $0.150 | $0.725 |
| arcee-ai | 1 | $0.250 | $0.800 |
| DeepSeek | 16 | $0.323 | $0.902 |
| ~deepseek | 3 | $0.128 | $0.926 |
| ~z-ai | 2 | $0.303 | $0.952 |
| Meituan | 1 | $0.300 | $1.20 |
| Baidu | 1 | $0.420 | $1.25 |
| MiniMax | 8 | $0.296 | $1.30 |
Recent Price Drops (83 models)
Models that have decreased in price since tracking began.
| Model | Provider | Previous Output $/M | Current Output $/M | Change |
|---|---|---|---|---|
| Ling-2.6-flash | inclusionai | $0.240 | $0.030 | -87.5% |
| MiMo-V2.5 | Xiaomi | $2.00 | $0.280 | -86.0% |
| GPT-5.6 Luna Pro | OpenAI | $6.00 | $1.20 | -80.0% |
| GPT-5.6 Luna | OpenAI | $6.00 | $1.20 | -80.0% |
| gpt-oss-120b (batch) | OpenAI | $0.600 | $0.136 | -77.3% |
| Ling-2.6-1T | inclusionai | $2.50 | $0.625 | -75.0% |
| MiMo-V2.5-Pro | Xiaomi | $3.00 | $0.870 | -71.0% |
| OpenAI GPT Latest | ~openai | $30.00 | $10.00 | -66.7% |
| GPT-5.6 Sol Pro | OpenAI | $30.00 | $10.00 | -66.7% |
| GPT-5.6 Sol | OpenAI | $30.00 | $10.00 | -66.7% |
| GPT-5.6 Sol Pro (batch) | OpenAI | $15.00 | $5.00 | -66.7% |
| GPT-5.6 Sol (batch) | OpenAI | $15.00 | $5.00 | -66.7% |
| DeepSeek V4 Flash 0423 | DeepSeek | $0.280 | $0.094 | -66.4% |
| Ling 3.0 Flash VL | inclusionai | $0.180 | $0.062 | -65.8% |
| DeepSeek V3.2 Speciale | DeepSeek | $1.20 | $0.431 | -64.1% |
| GLM 5.3 Flash (batch) | Zhipu AI | $0.500 | $0.200 | -60.0% |
| GLM Latest | ~z-ai | $4.40 | $1.76 | -59.9% |
| Grok 4.20 | xAI | $6.00 | $2.50 | -58.3% |
| Grok 4.20 Multi-Agent | xAI | $6.00 | $2.50 | -58.3% |
| GPT Luna Latest | ~openai | $1.20 | $0.500 | -58.3% |
| Llama Guard 3 8B | Meta | $0.060 | $0.030 | -50.0% |
| Gemini 3.6 Flash | $7.50 | $3.75 | -50.0% | |
| Gemini 3.6 Flash (batch) | $3.75 | $1.88 | -50.0% | |
| Mistral Large 3 2512 (batch) | Mistral AI | $1.50 | $0.750 | -50.0% |
| DeepSeek V4 Pro 0813 (batch) | DeepSeek | $3.96 | $1.98 | -50.0% |
Recent Price Increases (75 models)
| Model | Provider | Previous Output $/M | Current Output $/M | Change |
|---|---|---|---|---|
| Command R (08-2024) | Cohere | $0.600 | $10.00 | +1566.7% |
| Qwen3 30B A3B Thinking 2507 | Alibaba | $0.340 | $2.40 | +605.9% |
| Llama 3.2 11B Vision Instruct | Meta | $0.049 | $0.345 | +604.1% |
| Qwen2.5 Coder 32B Instruct | Alibaba | $0.200 | $1.00 | +400.0% |
| Qwen3 235B A22B Thinking 2507 | Alibaba | $0.600 | $2.30 | +283.3% |
| Qwen3 235B A22B Instruct 2507 | Alibaba | $0.100 | $0.350 | +250.0% |
| Llama 3 8B Instruct | Meta | $0.040 | $0.140 | +250.0% |
| gpt-oss-120b | OpenAI | $0.190 | $0.600 | +215.8% |
| MoonshotAI Kimi Latest | ~moonshotai | $3.49 | $10.53 | +201.9% |
| Solar Pro 4 | Upstage | $0.120 | $0.360 | +200.0% |
| Gemma 3 27B | $0.150 | $0.450 | +200.0% | |
| Gemma 3n 4B | $0.040 | $0.120 | +200.0% | |
| Hy3 preview | Tencent | $0.260 | $0.600 | +130.8% |
| Qwen3 VL 235B A22B Instruct | Alibaba | $0.880 | $1.90 | +115.9% |
| Qwen2.5 7B Instruct | Alibaba | $0.100 | $0.200 | +100.0% |
| MiniMax M3 (batch) | MiniMax | $0.600 | $1.20 | +100.0% |
| Inkling (batch) | thinkingmachines | $2.02 | $4.05 | +100.0% |
| Kimi K2.7 Code (batch) | Moonshot AI | $2.00 | $4.00 | +100.0% |
| Nemotron 3 Ultra (batch) | NVIDIA | $1.80 | $3.60 | +100.0% |
| Gemini 3.7 Flash | $1.88 | $3.75 | +100.0% | |
| Gemini 3.7 Flash (batch) | $0.938 | $1.88 | +100.0% | |
| Qwen3 30B A3B | Alibaba | $0.280 | $0.500 | +78.6% |
| DeepSeek V4 Flash Latest | ~deepseek | $0.180 | $0.320 | +77.8% |
| DeepSeek V4 Flash 0731 | DeepSeek | $0.180 | $0.320 | +77.8% |
| Granite 4.2 8B | IBM | $0.150 | $0.250 | +66.7% |
AI API Pricing Trends
AI model API pricing has been on a consistent downward trajectory since 2023. OpenAI's GPT-4 launched at $60/M output tokens; today, models with comparable capability cost under $5/M. This represents a 90%+ price reduction in under three years.
Key pricing trends observed across the industry:
- Race to the bottom: Competition between OpenAI, Anthropic, Google, and open-source providers drives continuous price cuts.
- Tiered pricing: Providers now offer multiple tiers (mini/flash models) at 10-50x lower cost than flagship models.
- Free tiers expanding: Google Gemini Flash, DeepSeek, and many open-source models available at zero cost through various providers.
- Caching discounts: Prompt caching can reduce input costs by 50-90% for repeated context.
Every Sunday at 23:00 UTC a cron job snapshots the full paid model catalog (currently 469 models across 68 providers) and writes a flat JSON file under data/weekly-snapshots/. The first snapshot in the table above was captured on 2026-02-16. Each row of every table on this page is sourced from those raw weekly files - no smoothing, no fills, no retroactive edits.
Most large providers adjust list pricing two to four times a year, typically alongside a new model generation. Mid-generation cuts of 50 to 90 percent are routine once a newer flagship ships at the same quality tier. The Recent Price Drops table above only counts changes greater than one percent, so cosmetic rounding does not show up as movement.
A handful of providers run experimental promotional pricing, particularly during model launches, and revert later. We treat every snapshot as authoritative for the week it was captured rather than smoothing those reversals away, so a temporary cut followed by a reversion shows as a drop in one week and an increase in another. This is intentional: it preserves the audit trail.
Google and DeepSeek have driven the steepest cuts in the dataset above, both by releasing flagship-tier reasoning models at fractions of incumbent pricing and by maintaining free tiers on Gemini Flash variants. Anthropic and OpenAI tend to cut older SKUs at the moment a new generation launches rather than adjusting the active flagship.
See /trackers/price-convergence-index. That page covers the LMC Price Convergence Index, which measures how tightly the frontier-tier price distribution clusters around its median in log space, with bias-corrected confidence intervals and a per-provider variance contribution breakdown. This page covers the per-model price trail: who moved when, and by how much.