驱动因素分析
分析驱动AI模型排名的关键因素。驱动因素代表最影响模型在排行榜中位置的特定信号 -- 包括正面助推和负面拖累。
独特驱动因素
6
有驱动因素的模型
300
主要正面驱动因素
Pricing
280 个模型
主要负面驱动因素
Capabilities
4 个模型
按频率排列的主要驱动因素
驱动因素频率
每个驱动因素在所有排名模型中出现的频率,按影响类型分类。
| 驱动因素 | 总计 | 净影响 |
|---|---|---|
| Capabilities | 297 | +240 |
| Pricing | 281 | +280 |
| Recency | 234 | +232 |
| Benchmarks | 225 | +196 |
| Context Window | 143 | +143 |
| Output Capacity | 20 | +20 |
正面驱动 - 有利因素
最常见的6个正面影响驱动因素,提升模型排名。
Pricing
280 个模型$25.00/M output tokens
Capabilities
244 个模型Supports reasoning, vision, tools, JSON mode, web search, streaming
Recency
232 个模型Released 2 months ago
Benchmarks
197 个模型Context Window
143 个模型1M token context window
Output Capacity
20 个模型Up to 128K output tokens per request
负面驱动 - 不利因素
最常见的2个负面影响驱动因素,拉低模型排名。
Capabilities
4 个模型Supports reasoning, vision, tools, JSON mode, web search, streaming
Benchmarks
1 个模型按模型分析驱动因素
按评分排列的前20个模型及其各自的驱动因素分析。
驱动因素信号分布
按底层信号类别分组的驱动因素,显示正面、负面和中性影响的分布。
| 信号 | 数量 |
|---|---|
| capability | 297 |
| pricing_tier | 281 |
| recency | 234 |
| benchmark | 225 |
| context_window | 143 |
| output_capacity | 20 |
方法论
驱动因素分析的工作原理。
驱动因素代表什么
驱动因素是最影响模型在排行榜中位置的特定因素。每个驱动因素捕获模型质量、定价、能力或市场表现的一个独特方面,这些方面共同构成综合排名评分。
如何计算
驱动因素来源于跨多个维度评估模型的评分算法。该算法识别出对每个模型最终排名影响最大的信号,然后将最大贡献者作为驱动因素呈现,附带相应的影响方向和指标值。
影响类型
正面 驱动因素帮助模型获得更高排名 -- 模型在此方面表现出色。 负面 驱动因素拉低模型排名 -- 这是模型的薄弱环节。 中性 驱动因素存在但不会显著影响排名方向。
Ranking drivers are the individual factors that push a model's composite score up or down. Positive drivers (like strong benchmark performance or competitive pricing) boost a model's rank, while negative drivers (like limited capabilities or high cost) pull it down.
Each driver represents the difference between a model's signal score and the average across all models, weighted by importance. Signals include capability breadth, pricing tier, context window size, recency, output capacity, and versatility.
Benchmark performance accounts for 90% of the score, from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations. Capabilities and context window serve as tiebreakers (10%), while output capacity and versatility each add 10%.