开源与专有AI模型
比较开源和专有模型在性能、价格、功能和稳定性方面的表现。追踪 300 个模型,帮助您决定哪种方法最适合您的需求。
快速对比
开源
公开可用的权重闭源
闭源 / 仅API开源与闭源
按评分排名的顶级开源模型
对比指标
| 指标 | 开源 | 闭源 |
|---|---|---|
| 模型数量 | 105 | 195 |
| 平均评分 | 60.3 | 71.8 |
| 中位分 | 64.1 | 76.9 |
| 最佳评分 | 86.7DeepSeek V4 Pro | 97.1Claude Fable 5 |
| 平均成本 ($/1M) | $0.946 | $12.31 |
| 免费模型 | 14 | 3 |
| 平均上下文窗口 | 340K | 634K |
| 稳定模型占比 | 3.8% | 4.6% |
| 脆弱模型占比 | 89.5% | 64.1% |
顶级开源模型
Top 20| # | 模型 | 评分 |
|---|---|---|
| 1 | DeepSeek V4 ProDeepSeek | 87 |
| 2 | DeepSeek V3.2DeepSeek | 81 |
| 3 | Gemma 4 31BGoogle | 81 |
| 4 | Gemma 4 31B (free)Google | 81 |
| 5 | Qwen3.5 397B A17BAlibaba | 79 |
| 6 | R1 0528DeepSeek | 79 |
| 7 | GLM 5.2Zhipu AI | 79 |
| 8 | GLM 5.2 (batch)Zhipu AI | 78 |
| 9 | GLM 5.1Zhipu AI | 78 |
| 10 | MiniMax M2.7MiniMax | 78 |
| 11 | MiniMax M2.5MiniMax | 78 |
| 12 | GLM 5Zhipu AI | 78 |
| 13 | Qwen3.5-122B-A10BAlibaba | 78 |
| 14 | Gemma 2 27BGoogle | 77 |
| 15 | Qwen3.5-27BAlibaba | 77 |
| 16 | MiMo-V2.5-ProXiaomi | 76 |
| 17 | Qwen3.5-35B-A3BAlibaba | 76 |
| 18 | Kimi K2.6Moonshot AI | 76 |
| 19 | GLM 4.7Zhipu AI | 75 |
| 20 | GLM 4.6Zhipu AI | 75 |
顶级专有模型
Top 20| # | 模型 | 评分 |
|---|---|---|
| 1 | Claude Fable 5Anthropic | 97 |
| 2 | Claude Fable 5 (batch)Anthropic | 97 |
| 3 | Claude Opus 5 (Fast)Anthropic | 95 |
| 4 | Claude Opus 5Anthropic | 95 |
| 5 | Claude Opus 4.8 (Fast)Anthropic | 95 |
| 6 | Claude Opus 4.8Anthropic | 95 |
| 7 | Claude Opus 4.7 (Fast)Anthropic | 95 |
| 8 | Claude Opus 4.7Anthropic | 95 |
| 9 | Claude Opus 4.7 (batch)Anthropic | 95 |
| 10 | Claude Opus 4.8 (batch)Anthropic | 95 |
| 11 | GPT-5.5 ProOpenAI | 93 |
| 12 | GPT-5.5 Pro (batch)OpenAI | 93 |
| 13 | GPT-5.5OpenAI | 93 |
| 14 | GPT-5.5 (batch)OpenAI | 93 |
| 15 | Gemini 3.1 Pro Preview Custom ToolsGoogle | 92 |
| 16 | Gemini 3.1 Pro PreviewGoogle | 92 |
| 17 | Gemini 3.1 Pro Preview (batch)Google | 92 |
| 18 | GPT-5.4 ProOpenAI | 92 |
| 19 | GPT-5.4 Pro (batch)OpenAI | 92 |
| 20 | GPT-5.4OpenAI | 92 |
功能对比
各组功能采用率| 功能 | 开源 | 闭源 |
|---|---|---|
| 视觉 | 33 (31.4%) | 160 (82.1%) |
| 函数调用 | 99 (94.3%) | 187 (95.9%) |
| 流式输出 | 105 (100.0%) | 195 (100.0%) |
| JSON模式 | 84 (80.0%) | 184 (94.4%) |
| 推理 | 85 (81.0%) | 156 (80.0%) |
| 网页搜索 | 0 (0.0%) | 137 (70.3%) |
| 图像输出 | 0 (0.0%) | 0 (0.0%) |
价格对比
开源模型价格
专有模型价格
结论
开源 领先于 free model availability, lower average pricing. 拥有 14 个免费模型,开源提供了最便捷的实验和原型开发入口。
闭源 领先于 average score, median score, model count, context window size, top model performance, capability coverage. 顶级专有模型 (Claude Fable 5) 达到 97 分,设定了当前的性能上限。
在 300 个追踪模型中(105 个开源,195 个专有),格局正在快速演变。开源模型在自托管、微调和成本控制方面表现出色,而专有模型通常在原始性能和托管API便利性方面领先。
The gap is narrowing rapidly. Open-source models like DeepSeek, Qwen, and LLaMA now compete with proprietary models on many benchmarks. However, proprietary models often still lead in raw performance on the most demanding tasks.
Open-source models offer full transparency, self-hosting capability, fine-tuning freedom, no vendor lock-in, and often lower costs. They are ideal for privacy-sensitive applications and organizations that need full control over their AI stack.
The top-scoring open-source model is shown in our leaderboard above. Rankings update hourly based on composite scores that combine benchmarks, pricing, capabilities, and community adoption.
We classify models based on whether their weights are publicly available for download and modification. Models with open weights but restrictive licenses are still counted as open source for this comparison.