Skip to content

性能衰退追踪器

检测AI模型可能出现的性能下降。此追踪器监控排名变动(排行榜上的位置变化),并标记在24小时或7天内显著下降的模型。"衰退分"越高,意味着越多的警告信号。

存在风险的模型

225

下降中 (7天)

212

不稳定

219

持续下降

1

按性能衰退风险评分排名的模型

LMMarketCap.com

存在风险的模型

225 个模型显示性能衰退迹象,按风险评分排名。更高的风险评分表示更令人担忧的性能趋势。

衰退分模型质量24小时排名7天排名严重程度
135Inkling Smallthinkingmachines71.6-10-60high
135Seed 1.6 FlashByteDance40.00-65high
135Seed 1.6ByteDance40.00-65high
135Nemotron 3 Nano 30B A3BNVIDIA40.00-65high
135Nemotron 3 Nano 30B A3B (free)NVIDIA40.00-65high
133gpt-oss-20bOpenAI57.40-64high
133gpt-oss-20b (free)OpenAI57.40-64high
133GPT-4o-mini (2024-07-18)OpenAI56.50-64high
133Qwen3 235B A22BAlibaba54.00-64high
133GPT-4.1 MiniOpenAI53.60-64high
133Kimi K2 ThinkingMoonshot AI53.30-64high
133DeepSeek V4 Flash Latest~deepseek40.00-64high
133DeepSeek V4 Flash 0731DeepSeek40.00-64high
133Qwen3.7 FlashAlibaba40.00-64high
133Nemotron 3 Ultra (free)NVIDIA40.00-64high
133Step 3.7 FlashStepFun40.00-64high
131Qwen3 30B A3BAlibaba64.10-63high
131o3 Mini HighOpenAI64.10-63high
131Mercury 2Inception61.00-63high
131Qwen3 8BAlibaba61.00-63high
131Nova 2 LiteAmazon60.50-63high
131Phi 4Microsoft60.20-63high
131Granite 4.1 8BIBM55.30-63high
131Olmo 3 32B ThinkAllen AI54.90-63high
131Llama 4 ScoutMeta54.70-63high
131Kimi K2.7 CodeMoonshot AI54.10-63high
131Kimi K2 0905Moonshot AI52.40-63high
131Kimi K2 0711Moonshot AI51.40-63high
131Claude 3 HaikuAnthropic51.30-63high
131Command ACohere50.80-63high
131Command R+ (08-2024)Cohere48.70-63high
131GPT-5 NanoOpenAI46.90-63high
131Llama 3.1 8B InstructMeta44.50-63high
131GPT-4.1 NanoOpenAI42.10-63high
131R1 Distill Llama 70BDeepSeek40.70-63high
131gpt-oss-120bOpenAI40.40-63high
131Laguna S 2.1poolside40.00-63high
131Laguna S 2.1 (free)poolside40.00-63high
131LongCat 2.0Meituan40.00-63high
131Kimi K3Moonshot AI40.00-63high
131KAT-Coder-Air V2.5Kuaishou40.00-63high
131KAT-Coder-Pro V2.5Kuaishou40.00-63high
131Grok Latest~x-ai40.00-63high
131Aion-3.0-Miniaion-labs40.00-63high
131Aion-3.0aion-labs40.00-63high
131Laguna XS 2.1poolside40.00-63high
131Laguna XS 2.1 (free)poolside40.00-63high
131Fugu Ultrasakana40.00-63high
131North Mini Code (free)Cohere40.00-63high
131Claude Fable Latest~anthropic40.00-63high
131Nemotron 3.5 Content Safety (free)NVIDIA40.00-63high
131Nemotron 3 UltraNVIDIA40.00-63high
131Perceptron Mk1perceptron40.00-63high
131Ring-2.6-1Tinclusionai40.00-63high
131GPT Chat LatestOpenAI40.00-63high
131Nemotron 3 Nano Omni (free)NVIDIA40.00-63high
131Anthropic Claude Haiku Latest~anthropic40.00-63high
131OpenAI GPT Mini Latest~openai40.00-63high
131Google Gemini Pro Latest~google40.00-63high
131MoonshotAI Kimi Latest~moonshotai40.00-63high
131Google Gemini Flash Latest~google40.00-63high
131Anthropic Claude Sonnet Latest~anthropic40.00-63high
131OpenAI GPT Latest~openai40.00-63high
131Ling-2.6-1Tinclusionai40.00-63high
131Ling-2.6-flashinclusionai40.00-63high
131Claude Opus Latest~anthropic40.00-63high
131Lyria 3 Pro PreviewGoogle40.00-63high
131Lyria 3 Clip PreviewGoogle40.00-63high
131KAT-Coder-Pro V2Kuaishou40.00-63high
131Reka Edgerekaai40.00-63high
131Mistral Small 4Mistral AI40.00-63high
131Nemotron 3 SuperNVIDIA40.00-63high
131Nemotron 3 Super (free)NVIDIA40.00-63high
131Seed-2.0-LiteByteDance40.00-63high
131Seed-2.0-MiniByteDance40.00-63high
131Aion-2.0aion-labs40.00-63high
129Trinity Large Thinkingarcee-ai63.90-62high
129Kimi K2.5Moonshot AI59.10-62high
129Command R (08-2024)Cohere48.7+1-62high
129Qwen3.6 FlashAlibaba40.00-62high
129Qwen3.6 35B A3BAlibaba40.00-62high
129Qwen3.6 27BAlibaba40.00-62high
129Qwen3 Coder NextAlibaba40.00-62high
129Solar Pro 3Upstage40.00-62high
129Palmyra X5Writer40.00-62high
129GPT AudioOpenAI40.00-62high
129GPT Audio MiniOpenAI40.00-62high
127Hy3Tencent74.00-61high
127GPT-4o (2024-08-06)OpenAI71.20-61high
127GPT-4oOpenAI71.20-61high
127GPT-4o (2024-05-13)OpenAI71.20-61high
127GPT-4OpenAI64.80-61high
127Qwen3 235B A22B Instruct 2507Alibaba64.70-61high
127GPT-5 MiniOpenAI63.90-61high
127GLM 4.7 FlashZhipu AI63.70-61high
127Mixtral 8x22B InstructMistral AI63.40-61high
127GLM 4.5VZhipu AI62.50-61high
125DeepSeek V3 0324DeepSeek71.8+1-60high
125Mistral Medium 3.5Mistral AI71.7+1-60high
125MiniMax M1MiniMax70.80-60high
125GLM 4.5 AirZhipu AI70.70-60high
125Mistral LargeMistral AI65.90-60high
125Composer 2Cursor65.70-60high
125Composer 2 FastCursor65.70-60high
125GLM 4.6VZhipu AI65.50-60high
125Qwen3 235B A22B Thinking 2507Alibaba65.50-60high
125Llama 3.1 70B InstructMeta65.30-60high
123Hy3 previewTencent68.20-59high
123Qwen3 Next 80B A3B InstructAlibaba66.90-59high
123Llama 3.3 70B InstructMeta66.80-59high
123GPT-4 TurboOpenAI66.70-59high
123Qwen3.5-9BAlibaba66.50-59high
123Step 3.5 FlashStepFun66.10-59high
123Qwen3 30B A3B Thinking 2507Alibaba64.10-59high
121Qwen3 Max ThinkingAlibaba68.20-58high
121Llama 4 MaverickMeta67.90-58high
121GPT-4.1OpenAI67.70-58high
121Qwen3 MaxAlibaba67.40-58high
121Mistral Large 3 2512Mistral AI67.00-58high
119MiniMax M2MiniMax72.0+1-57high
119Qwen3 VL 235B A22B InstructAlibaba69.40-57high
119DeepSeek V3.1 TerminusDeepSeek69.30-57high
119GPT-4o-miniOpenAI69.30-57high
117MiMo-V2.5Xiaomi73.00-56high
117Gemma 4 26B A4B Google73.00-56high
117Gemma 4 26B A4B (free)Google73.00-56high
117GLM 5V TurboZhipu AI72.3+1-56high
117o4 Mini HighOpenAI72.1+1-56high
117Claude Haiku 4.5Anthropic69.50-56high
117DeepSeek V3DeepSeek69.50-56high
117MiniMax M2-herMiniMax69.10-56high
117Qwen3.5-FlashAlibaba68.60-56high
115Qwen3.6 PlusAlibaba74.10-55high
115o1-proOpenAI73.60-55high
113Inklingthinkingmachines74.00-54high
109R1DeepSeek74.20-52high
107Qwen3.6 Max PreviewAlibaba74.80-51high
107MiniMax M3MiniMax74.30-51high
105Qwen3.7 PlusAlibaba75.90-50high
105Claude Sonnet 4Anthropic74.40-50high
105o1OpenAI74.40-50high
103GLM 4.5Zhipu AI75.10-49high
101MiniMax M2.5MiniMax78.00-48high
101GLM 5Zhipu AI78.00-48high
101Kimi K2.6Moonshot AI75.80-48high
101DeepSeek V3.2 ExpDeepSeek71.8+1-48high
99MiMo-V2.5-ProXiaomi76.20-47high
99Qwen3.5-35B-A3BAlibaba76.00-47high
99o3 MiniOpenAI75.30-47high
99Qwen3 VL 235B A22B ThinkingAlibaba69.40-47high
97Qwen3.5-122B-A10BAlibaba77.70-46high
97Gemma 2 27BGoogle77.40-46high
95Qwen3.5-27BAlibaba77.00-45high
95Gemini 3.5 Flash LiteGoogle76.90-45high
95DeepSeek V3.1DeepSeek71.8+1-45high
95GPT-4 Turbo PreviewOpenAI64.80-45high
93GLM 5.2Zhipu AI78.50-44high
93MiniMax M2.1MiniMax72.0+1-44high
91Gemini 3.5 FlashGoogle79.00-43high
89Gemini 2.5 FlashGoogle79.10-42high
87Gemini 3.1 Flash Lite PreviewGoogle79.30-41high
87Gemini 2.5 Flash LiteGoogle79.10-41high
87Qwen3 Next 80B A3B ThinkingAlibaba66.90-41high
85Muse Spark 1.1meta80.20-40high
85GPT-5.4 MiniOpenAI79.30-40high
83Qwen3.5 397B A17BAlibaba79.40-39high
83R1 0528DeepSeek79.40-39high
83GPT-5.4 NanoOpenAI79.30-39high
81Gemini 3.6 FlashGoogle80.00-38high
79DeepSeek V3.2DeepSeek81.30-37high
77o4 MiniOpenAI81.40-36high
77Gemma 4 31BGoogle80.50-36high
77Gemma 4 31B (free)Google80.50-36high
75Claude Opus 4Anthropic82.10-35high
71Gemini 2.5 Pro Preview 06-05Google83.50-33high
71Gemini 2.5 Pro Preview 05-06Google83.50-33high
71Claude Sonnet 4.5Anthropic82.40-33high
71GLM 4.7Zhipu AI75.10-33high
71Mistral Large 2407Mistral AI65.90-33high
69Gemini 2.5 ProGoogle83.50-32high
69GLM 5.1Zhipu AI78.00-32high
67Claude Opus 4.5Anthropic85.10-31high
65Claude Sonnet 4.6Anthropic85.20-30high
63Gemini 3 Flash PreviewGoogle88.40-29high
61GPT-5OpenAI88.70-28high
61GPT-5.1-Codex-MiniOpenAI87.80-28high
61o3OpenAI86.70-28high
59DeepSeek V4 ProDeepSeek86.70-27high
59o3 ProOpenAI86.70-27high
57GPT-5 ProOpenAI88.70-26high
47GPT-5.6 SolOpenAI89.00-21high
47GLM 4.6Zhipu AI75.10-21high
45GPT-5.6 Sol ProOpenAI89.00-20high
43GPT-5.6 TerraOpenAI89.00-19high
43GPT-5.1-Codex-MaxOpenAI88.70-19high
43GPT-5.1OpenAI88.70-19high
43GPT-5.1-CodexOpenAI88.70-19high
41GPT-5.6 Terra ProOpenAI89.00-18high
39GPT-5.6 LunaOpenAI89.00-17high
37GPT-5.6 Luna ProOpenAI89.00-16high
35Claude Opus 4.6Anthropic90.40-15high
33GPT-5.2OpenAI90.50-14high
31GPT-5.2 ProOpenAI90.50-13high
29GPT-5.2-CodexOpenAI90.50-12high
25GPT-5.4OpenAI91.90-10high
23GPT-5.4 ProOpenAI91.90-9high
21Gemini 3.1 Pro Preview Custom ToolsGoogle92.20-8high
21Gemini 3.1 Pro PreviewGoogle92.20-8high
21GLM 5 TurboZhipu AI78.00-8high
19GPT-5.5OpenAI92.70-7high
10Claude Opus 4.7 (Fast)Anthropic95.10-5high
10Claude Opus 4.7Anthropic95.10-5high
5Claude Opus 5 (Fast)Anthropic95.10+171medium
5Claude Opus 5Anthropic95.10+171medium
5GPT-5.3 ChatOpenAI90.50+10medium
5GPT-5.2 ChatOpenAI90.50+40medium
5Claude Sonnet 5Anthropic85.20+123medium
5Qwen3.7 MaxAlibaba74.80+71medium
5Qwen3.5 Plus 2026-04-20Alibaba74.80+83medium
5Qwen3.5 Plus 2026-02-15Alibaba68.20+55medium
5Qwen3 30B A3B Instruct 2507Alibaba64.10+65medium
4Claude Opus 4.1Anthropic82.10-2low
2Claude Opus 4.8 (Fast)Anthropic95.10-1low
2Claude Opus 4.8Anthropic95.10-1low
2MiniMax M2.7MiniMax78.00-1low

稳定模型

7 个模型无下降且排名状态稳定。这些模型表现一致。

#模型评分24h7d状态
1Claude Fable 5Anthropic97.100stable
11GPT-5.5 ProOpenAI92.700stable
23GPT-5.3-CodexOpenAI90.500stable
155GPT-4o (2024-11-20)OpenAI71.20+4stable
294Falcon-H1-Arabic 34B InstructTII40.00+1stable
295Falcon-H1-Arabic 7B InstructTII40.00+1stable
296Falcon-H1-Arabic 3B InstructTII40.00+1stable

如何检测性能衰退

我们的性能衰退检测系统使用多种信号来识别可能正在下降的模型。

下降中 (7天)

7天排名变化超过-2位的模型。一周内持续下降超过两个名次,表明该模型可能正在被竞争对手超越或出现性能问题。

脆弱状态

被评分系统标记为"脆弱"的模型。这些模型的性能指标不一致或评分处于边界,评估数据的微小变化可能导致显著波动。

持续下降

在24小时和7天两个时间维度上均下降的模型。当模型在短期和中期窗口都在失去排名时,表明这是持续的下降趋势而非暂时波动。

风险评分

性能衰退风险评分综合了多个信号:7天排名下降权重2倍,24小时排名下降权重1倍,脆弱状态额外加5分。更高的评分表示有更大的性能衰退风险。

相关

Frequently Asked Questions

The tracker uses a multi-signal approach: it monitors 7-day rank decline (weighted 2x), 24-hour rank drops (weighted 1x), and fragile state classification (+5 points). Models are scored on a degradation risk scale where higher values indicate more warning signs of performance decline.

A fragile state indicates that a model has inconsistent performance metrics or borderline scores that could shift significantly with small changes in evaluation data. Fragile models are at higher risk of further ranking drops and warrant closer monitoring.

Yes, models can recover. Degradation may be temporary due to API issues, benchmark fluctuations, or scoring recalibrations. Models that show sustained decline over multiple weeks are more concerning than those with short-term dips. The tracker monitors both 24-hour and 7-day windows to help distinguish temporary noise from real trends.

AI Model Degradation Tracker - Detect Performance Drops | LM Market Cap