Skip to content

Small Language Models

96 lightweight AI models under $1/1M tokens. Small language models (SLMs) are optimized for speed, low cost, and edge deployment - ideal for mobile apps, IoT, chatbots, and high-volume production workloads.

96
Under $1/1M
13
Free
61
Open Source
87
+ Tool Use

Small Language Models - Ranked by Efficiency

#ModelScore
1GPT-5.6 Luna Pro (batch)OpenAI89
2GPT-5.6 Luna (batch)OpenAI89
3Gemma 4 31B (free)Google81
4DeepSeek V3.2DeepSeek83
5Gemma 4 31BGoogle81
6Gemini 2.5 Flash Lite (batch)Google79
7GLM 5.3 Flash (batch)Zhipu AI78
8GLM 5.2 (free)Zhipu AI76
9Gemini 2.5 Flash LiteGoogle79
10GPT-5 Nano (batch)OpenAI77
11GPT-5.4 Nano (batch)OpenAI79
12GLM 5.3 FlashZhipu AI78
13Gemma 4 26B A4B (free)Google73
14Qwen3.8 27B (free)Alibaba73
15Inkling (free)thinkingmachines73
16Hy3Tencent74
17Gemma 2 27BGoogle77
18MiMo-V2.5Xiaomi73
19Gemma 4 26B A4B Google73
20Inkling Small (free)thinkingmachines69
21DeepSeek V3.2 ExpDeepSeek72
22MiMo-V2.6-ProXiaomi76
23MiMo-V2.5-ProXiaomi76
24Qwen3.5-FlashAlibaba69
25Qwen3.5-9BAlibaba67
26Llama 4 MaverickMeta71
27Llama 3.3 70B InstructMeta67
28GPT-4o-miniOpenAI69
29Step 3.5 FlashStepFun66
30Hy3 previewTencent68

When to Use Small Language Models

High-Volume Applications

Processing millions of requests per day? SLMs cost 10-100x less than premium models. A chatbot handling 1M messages/month costs ~$100 with budget models vs $10,000+ with premium ones.

Edge & Mobile Deployment

Open-source SLMs can run on consumer hardware - laptops, phones, or edge devices. Models like Phi, Gemma, and small Llama variants fit in 4-8GB of RAM.

Low-Latency Requirements

Smaller models respond faster. For real-time applications like autocomplete, classification, or chat, SLMs deliver sub-100ms responses.

Task-Specific Workloads

Many tasks - classification, extraction, summarization, translation - don't need the largest models. A well-chosen SLM can match premium model quality on focused tasks.

Frequently Asked Questions

Small language models are AI models with fewer parameters, typically under 10 billion. They run faster, cost less, and can operate on edge devices while still handling many common tasks like text generation, summarization, and simple coding assistance.

Use SLMs when you need low latency, low cost, or on-device deployment. Use full LLMs when you need complex reasoning, creative writing, or state-of-the-art accuracy. SLMs are ideal for chatbots, simple Q&A, and high-volume applications where cost matters more than peak performance.

SLMs trade some capability for speed and efficiency. Modern SLMs like Phi-3 and Gemma 2 can match older large models on many benchmarks. For specialized tasks, a fine-tuned SLM can outperform a general-purpose LLM while being 10-100x cheaper to run.

Best Small Language Models - SLMs Ranked (2026) | LM Market Cap