功能覆盖追踪器
分析300个追踪模型中哪些AI功能最常见和最稀有,以及每项功能如何与综合评分相关。发现采用趋势、评分溢价和功能组合。
概览
功能采用率 (%)
功能采用
按采用率排序| 功能 | 模型 | 采用率 | 评分溢价 |
|---|---|---|---|
Streaming | 300 | 100.0% | +67.8 |
Function Calling | 286 | 95.3% | +12.0 |
JSON Mode | 268 | 89.3% | +16.3 |
Reasoning | 241 | 80.3% | +11.5 |
Vision | 193 | 64.3% | +15.2 |
Web Search | 137 | 45.7% | +19.7 |
Image Output | 0 | 0.0% | -67.8 |
评分溢价分析
每项功能与高评分的关联程度功能组合
最受欢迎的多功能组合| # | 组合 | 模型 | 平均评分 |
|---|---|---|---|
| 1 | Function Calling + Streaming | 286 | 68.3 |
| 2 | JSON Mode + Streaming | 268 | 69.5 |
| 3 | Function Calling + JSON Mode | 258 | 69.9 |
| 4 | Function Calling + JSON Mode + Streaming | 258 | 69.9 |
| 5 | Reasoning + Streaming | 241 | 70.0 |
| 6 | Function Calling + Reasoning | 234 | 70.4 |
| 7 | Function Calling + Reasoning + Streaming | 234 | 70.4 |
| 8 | JSON Mode + Reasoning | 213 | 72.2 |
| 9 | JSON Mode + Reasoning + Streaming | 213 | 72.2 |
| 10 | Function Calling + JSON Mode + Reasoning | 208 | 72.4 |
"全栈"模型
拥有全部7项功能的模型目前没有模型同时具备全部7项功能(视觉、函数调用、流式输出、JSON模式、推理、网页搜索和图像输出)。"全栈"模型是指支持所有追踪功能的模型。
各服务商功能对比
拥有3个以上模型的服务商| 提供商 | Vision | Function Calling | Streaming | JSON Mode | Reasoning | Web Search | Image Output |
|---|---|---|---|---|---|---|---|
| OpenAI(86) | 75/8687% | 84/8698% | 86/86100% | 86/86100% | 63/8673% | 72/8684% | 00% |
| Alibaba(31) | 17/3155% | 31/31100% | 31/31100% | 31/31100% | 25/3181% | 00% | 00% |
| Anthropic(28) | 28/28100% | 28/28100% | 28/28100% | 24/2886% | 27/2896% | 28/28100% | 00% |
| Google(27) | 26/2796% | 24/2789% | 27/27100% | 27/27100% | 24/2789% | 20/2774% | 00% |
| Zhipu AI(13) | 3/1323% | 13/13100% | 13/13100% | 12/1392% | 13/13100% | 00% | 00% |
| DeepSeek(12) | 00% | 11/1292% | 12/12100% | 11/1292% | 10/1283% | 00% | 00% |
| NVIDIA(9) | 2/922% | 8/989% | 9/9100% | 5/956% | 9/9100% | 00% | 00% |
| Moonshot AI(8) | 5/863% | 8/8100% | 8/8100% | 7/888% | 6/875% | 00% | 00% |
| MiniMax(8) | 2/825% | 7/888% | 8/8100% | 6/875% | 7/888% | 00% | 00% |
| Mistral AI(6) | 3/650% | 6/6100% | 6/6100% | 6/6100% | 2/633% | 00% | 00% |
| inclusionai(5) | 00% | 5/5100% | 5/5100% | 3/560% | 3/560% | 00% | 00% |
| xAI(5) | 5/5100% | 4/580% | 5/5100% | 5/5100% | 5/5100% | 5/5100% | 00% |
| Meta(5) | 2/540% | 5/5100% | 5/5100% | 5/5100% | 00% | 00% | 00% |
| poolside(4) | 00% | 4/4100% | 4/4100% | 00% | 4/4100% | 00% | 00% |
| Cohere(4) | 00% | 3/475% | 4/4100% | 3/475% | 1/425% | 00% | 00% |
| ~anthropic(4) | 4/4100% | 4/4100% | 4/4100% | 4/4100% | 4/4100% | 4/4100% | 00% |
| ByteDance(4) | 4/4100% | 4/4100% | 4/4100% | 4/4100% | 4/4100% | 00% | 00% |
| thinkingmachines(3) | 3/3100% | 3/3100% | 3/3100% | 1/333% | 3/3100% | 00% | 00% |
| Kuaishou(3) | 00% | 3/3100% | 3/3100% | 3/3100% | 00% | 00% | 00% |
| aion-labs(3) | 00% | 3/3100% | 3/3100% | 3/3100% | 3/3100% | 00% | 00% |
| TII(3) | 00% | 3/3100% | 3/3100% | 00% | 3/3100% | 00% | 00% |
We track seven key capabilities: Vision (image understanding), Function Calling (tool use), Streaming (real-time output), JSON Mode (structured output), Reasoning (chain-of-thought), Web Search (live information retrieval), and Image Output (image generation).
Score premium measures how much higher (or lower) the average composite score is for models that have a specific capability compared to those that lack it. A positive premium means models with that capability tend to score higher overall.
A full stack model supports all seven tracked capabilities: vision, function calling, streaming, JSON mode, reasoning, web search, and image output. These are the most versatile models available.
The most and least common capabilities are shown in the Overview section above. Adoption rates vary widely, with some capabilities like streaming being near-universal while others like image output are much rarer.