Skip to content

Small Language Models

89 lightweight AI models under $1/1M tokens. Small language models (SLMs) are optimized for speed, low cost, and edge deployment - ideal for mobile apps, IoT, chatbots, and high-volume production workloads.

89
Under $1/1M
17
Free
59
Open Source
82
+ Tool Use

Small Language Models - Ranked by Efficiency

#ModelScore
1GPT-5.6 Luna ProOpenAI89
2GPT-5.6 Luna Pro (batch)OpenAI89
3GPT-5.6 LunaOpenAI89
4GPT-5.6 Luna (batch)OpenAI89
5Gemma 4 31B (free)Google81
6DeepSeek V4 ProDeepSeek87
7DeepSeek V3.2DeepSeek81
8Gemma 4 31BGoogle81
9Gemini 2.5 Flash Lite (batch)Google79
10Gemini 2.5 Flash LiteGoogle79
11GPT-5 Nano (batch)OpenAI77
12GPT-5.4 Nano (batch)OpenAI79
13Gemma 4 26B A4B (free)Google73
14GLM 5.2Zhipu AI79
15Gemma 2 27BGoogle77
16MiMo-V2.5Xiaomi73
17Gemma 4 26B A4B Google73
18MiniMax M2.5MiniMax78
19Hy3Tencent74
20MiniMax M3 (batch)MiniMax74
21DeepSeek V3.2 ExpDeepSeek72
22MiMo-V2.5-ProXiaomi76
23Hy3 previewTencent68
24Qwen3.5-FlashAlibaba69
25Qwen3.5-9BAlibaba67
26Llama 3.3 70B InstructMeta67
27GPT-4o-miniOpenAI69
28Step 3.5 FlashStepFun66
29DeepSeek V3.1DeepSeek72
30GLM 4.5 AirZhipu AI71

When to Use Small Language Models

High-Volume Applications

Processing millions of requests per day? SLMs cost 10-100x less than premium models. A chatbot handling 1M messages/month costs ~$100 with budget models vs $10,000+ with premium ones.

Edge & Mobile Deployment

Open-source SLMs can run on consumer hardware - laptops, phones, or edge devices. Models like Phi, Gemma, and small Llama variants fit in 4-8GB of RAM.

Low-Latency Requirements

Smaller models respond faster. For real-time applications like autocomplete, classification, or chat, SLMs deliver sub-100ms responses.

Task-Specific Workloads

Many tasks - classification, extraction, summarization, translation - don't need the largest models. A well-chosen SLM can match premium model quality on focused tasks.

Frequently Asked Questions

Small language models are AI models with fewer parameters, typically under 10 billion. They run faster, cost less, and can operate on edge devices while still handling many common tasks like text generation, summarization, and simple coding assistance.

Use SLMs when you need low latency, low cost, or on-device deployment. Use full LLMs when you need complex reasoning, creative writing, or state-of-the-art accuracy. SLMs are ideal for chatbots, simple Q&A, and high-volume applications where cost matters more than peak performance.

SLMs trade some capability for speed and efficiency. Modern SLMs like Phi-3 and Gemma 2 can match older large models on many benchmarks. For specialized tasks, a fine-tuned SLM can outperform a general-purpose LLM while being 10-100x cheaper to run.

Best Small Language Models - SLMs Ranked (2026) | LM Market Cap