Skip to content

Self-Hosted AI Models

153 open-source AI models you can run on your own infrastructure. Self-hosting gives you complete data privacy, zero per-token costs, and full control over the model and its behavior.

153
Open Source
31
Providers
125
+ Tool Use
52
+ Vision
98
+ Reasoning

Top Open Source Providers

Alibaba
32 models
DeepSeek
12 models
Zhipu AI
11 models
NVIDIA
11 models
Mistral AI
10 models
Google
9 models
Moonshot AI
8 models
Meta
8 models

All Self-Hostable Models - Ranked by Score

#ModelScore
1DeepSeek V4 ProDeepSeek87
2DeepSeek V3.2DeepSeek81
3Gemma 4 31BGoogle81
4Gemma 4 31B (free)Google81
5Qwen3.5 397B A17BAlibaba79
6R1 0528DeepSeek79
7GLM 5.2Zhipu AI79
8GLM 5.2 (batch)Zhipu AI78
9GLM 5.1Zhipu AI78
10MiniMax M2.7MiniMax78
11MiniMax M2.5MiniMax78
12GLM 5Zhipu AI78
13Qwen3.5-122B-A10BAlibaba78
14Gemma 2 27BGoogle77
15Qwen3.5-27BAlibaba77
16MiMo-V2.5-ProXiaomi76
17Qwen3.5-35B-A3BAlibaba76
18Kimi K2.6Moonshot AI76
19GLM 4.7Zhipu AI75
20GLM 4.6Zhipu AI75
21GLM 4.5Zhipu AI75
22MiniMax M3MiniMax74
23R1DeepSeek74
24MiniMax M3 (batch)MiniMax74
25Inklingthinkingmachines74
26Hy3Tencent74
27MiMo-V2.5Xiaomi73
28Gemma 4 26B A4B Google73
29Gemma 4 26B A4B (free)Google73
30Inkling (batch)thinkingmachines73
31MiniMax M2.1MiniMax72
32MiniMax M2MiniMax72
33DeepSeek V3.2 ExpDeepSeek72
34DeepSeek V3.1DeepSeek72
35DeepSeek V3 0324DeepSeek72
36Inkling Smallthinkingmachines72
37GLM 4.5 AirZhipu AI71
38DeepSeek V3DeepSeek70
39Qwen3 VL 235B A22B ThinkingAlibaba69
40Qwen3 VL 235B A22B InstructAlibaba69

Self-Hosting AI Models

Complete Data Privacy

Your data never leaves your infrastructure. Critical for healthcare, finance, legal, and government use cases where data residency and privacy regulations apply.

Zero Per-Token Cost

After the initial hardware investment, there are no per-request charges. At high volumes, self-hosting can be 10-100x cheaper than API-based services.

Popular Hosting Options

Run models with vLLM, Ollama, text-generation-inference, or llama.cpp. Most can run on consumer GPUs (RTX 4090) for smaller models, or cloud GPUs (A100, H100) for larger ones.

Fine-Tuning & Customization

Self-hosted models can be fine-tuned on your own data, creating domain-specific versions that outperform general-purpose models for your use case.

Frequently Asked Questions

Self-hosting gives you complete data privacy (no data leaves your servers), eliminates per-token API costs, removes rate limits, enables offline operation, and allows fine-tuning for your specific use case.

Requirements depend on model size. Small models (7B parameters) run on consumer GPUs with 8GB VRAM. Medium models (13-30B) need 24GB+ VRAM. Large models (70B+) require multiple high-end GPUs or specialized inference hardware.

Popular tools include Ollama (easiest setup), llama.cpp (most efficient for CPU inference), vLLM (fastest for GPU serving), and text-generation-webui (feature-rich UI). Each excels at different use cases.

For high-volume usage (thousands of requests/day), self-hosting is significantly cheaper. For low-volume or sporadic use, API access is more cost-effective since you avoid hardware costs and maintenance overhead.

Best Self-Hosted AI Models - Run LLMs Locally (2026) | LM Market Cap