Skip to content

Self-Hosted AI Models

177 open-source AI models you can run on your own infrastructure. Self-hosting gives you complete data privacy, zero per-token costs, and full control over the model and its behavior.

177
Open Source
33
Providers
145
+ Tool Use
68
+ Vision
118
+ Reasoning

Top Open Source Providers

Alibaba
36 models
DeepSeek
16 models
Zhipu AI
15 models
Mistral AI
13 models
NVIDIA
10 models
Google
8 models
Moonshot AI
8 models
Meta
8 models

All Self-Hostable Models - Ranked by Score

#ModelScore
1DeepSeek V3.2DeepSeek83
2Gemma 4 31BGoogle81
3Gemma 4 31B (free)Google81
4Qwen3.5 397B A17BAlibaba79
5R1 0528DeepSeek79
6GLM 5.3Zhipu AI79
7GLM 5.3 (batch)Zhipu AI79
8GLM 5.2Zhipu AI79
9GLM 5.3 FlashZhipu AI78
10GLM 5.1Zhipu AI78
11GLM 5Zhipu AI78
12GLM 5.3 Flash (batch)Zhipu AI78
13Qwen3.5-122B-A10BAlibaba78
14Gemma 2 27BGoogle77
15Qwen3.5-27BAlibaba77
16MiMo-V2.6-ProXiaomi76
17MiMo-V2.5-ProXiaomi76
18Qwen3.5-35B-A3BAlibaba76
19GLM 5.2 (free)Zhipu AI76
20Kimi K2.6Moonshot AI76
21GLM 4.7Zhipu AI75
22GLM 4.6Zhipu AI75
23GLM 4.5Zhipu AI75
24Hy3Tencent74
25MiniMax M3MiniMax74
26R1DeepSeek74
27Qwen3.8 27BAlibaba73
28MiMo-V2.5Xiaomi73
29Gemma 4 26B A4B Google73
30Gemma 4 26B A4B (free)Google73
31Qwen3.8 27B (free)Alibaba73
32Inklingthinkingmachines73
33Inkling (free)thinkingmachines73
34DeepSeek V3.2 ExpDeepSeek72
35DeepSeek V3.1DeepSeek72
36DeepSeek V3 0324DeepSeek72
37MiniMax M2.7MiniMax72
38MiniMax M2.5MiniMax72
39MiniMax M2.1MiniMax71
40MiniMax M2MiniMax71

Self-Hosting AI Models

Complete Data Privacy

Your data never leaves your infrastructure. Critical for healthcare, finance, legal, and government use cases where data residency and privacy regulations apply.

Zero Per-Token Cost

After the initial hardware investment, there are no per-request charges. At high volumes, self-hosting can be 10-100x cheaper than API-based services.

Popular Hosting Options

Run models with vLLM, Ollama, text-generation-inference, or llama.cpp. Most can run on consumer GPUs (RTX 4090) for smaller models, or cloud GPUs (A100, H100) for larger ones.

Fine-Tuning & Customization

Self-hosted models can be fine-tuned on your own data, creating domain-specific versions that outperform general-purpose models for your use case.

Frequently Asked Questions

Self-hosting gives you complete data privacy (no data leaves your servers), eliminates per-token API costs, removes rate limits, enables offline operation, and allows fine-tuning for your specific use case.

Requirements depend on model size. Small models (7B parameters) run on consumer GPUs with 8GB VRAM. Medium models (13-30B) need 24GB+ VRAM. Large models (70B+) require multiple high-end GPUs or specialized inference hardware.

Popular tools include Ollama (easiest setup), llama.cpp (most efficient for CPU inference), vLLM (fastest for GPU serving), and text-generation-webui (feature-rich UI). Each excels at different use cases.

For high-volume usage (thousands of requests/day), self-hosting is significantly cheaper. For low-volume or sporadic use, API access is more cost-effective since you avoid hardware costs and maintenance overhead.

Best Self-Hosted AI Models - Run LLMs Locally (2026) | LM Market Cap