Skip to content

Change Significance Tracker

Not all rank changes are meaningful. Some are random noise. This page uses statistical analysis to tell you which model score movements are real trends vs. normal fluctuation, so you know which changes to pay attention to.

Models Analyzed

228

Significant Changes

7

Noise (Not Significant)

221

Both Timeframes

0

What This Means

  • 7 of 228 models have score changes that are statistically significant - these are real performance shifts, not random noise.
  • 221 models show score changes within their normal variation range - don't read too much into small rank shifts for these models.
  • How to use this: When a model's rank changes, check here first. If it's not flagged as significant, the change is likely temporary noise. If it is significant, the model is genuinely improving or declining.

Top Significant Changes (by |Z-Score|)

LMMarketCap.com

Real Performance Shifts

7 models whose recent scores deviate enough from their historical average to be considered a real change (not noise). Sorted by how extreme the change is. Z-Score measures how unusual the change is - values beyond ±1.96 mean there's a 95% chance the change is real.

ModelCurrentZ-ScoreDirection
Qwen3 30B A3B Instruct 2507Alibaba64.14.47Improvement
GLM 5.1Zhipu AI78.03.91Improvement
Qwen3.5 Plus 2026-04-20Alibaba74.83.74Improvement
Qwen3.7 MaxAlibaba74.83.32Improvement
GPT-4o (2024-11-20)OpenAI71.22.30Improvement
Claude Opus 4.7 (Fast)Anthropic95.12.24Improvement
Claude Sonnet 5Anthropic85.22.24Improvement

Short-Term vs. Sustained Changes

A model changing rank in 24 hours could be a blip. But if it's also moving over 7 days, that's a real trend. Models flagged on both timeframes are the most important to watch - they represent confirmed, sustained performance shifts.

Significant on Both Timeframes(strongest signals)

No models currently significant on both daily and weekly timeframes.

Daily Only(may be noise)

No models with daily-only significance.

Weekly Only(building trend)

ModelScore24h Change7d Change
Claude Opus 5 (Fast)Anthropic95.10+171
Claude Opus 5Anthropic95.10+171
GPT-5.5OpenAI92.70-7
Gemini 3.1 Pro Preview Custom ToolsGoogle92.20-8
Gemini 3.1 Pro PreviewGoogle92.20-8
GPT-5.4 ProOpenAI91.90-9
GPT-5.4OpenAI91.90-10
GPT-5.3 ChatOpenAI90.50+10
GPT-5.2-CodexOpenAI90.50-12
GPT-5.2 ChatOpenAI90.50+40
GPT-5.2 ProOpenAI90.50-13
GPT-5.2OpenAI90.50-14
Claude Opus 4.6Anthropic90.40-15
GPT-5.6 Luna ProOpenAI89.00-16
GPT-5.6 LunaOpenAI89.00-17
GPT-5.6 Terra ProOpenAI89.00-18
GPT-5.6 TerraOpenAI89.00-19
GPT-5.6 Sol ProOpenAI89.00-20
GPT-5.6 SolOpenAI89.00-21
GPT-5.1-Codex-MaxOpenAI88.70-19
GPT-5.1OpenAI88.70-19
GPT-5.1-CodexOpenAI88.70-19
GPT-5 ProOpenAI88.70-26
GPT-5OpenAI88.70-28
Gemini 3 Flash PreviewGoogle88.40-29
GPT-5.1-Codex-MiniOpenAI87.80-28
DeepSeek V4 ProDeepSeek86.70-27
o3 ProOpenAI86.70-27
o3OpenAI86.70-28
Claude Sonnet 5Anthropic85.20+123
Claude Sonnet 4.6Anthropic85.20-30
Claude Opus 4.5Anthropic85.10-31
Gemini 2.5 ProGoogle83.50-32
Gemini 2.5 Pro Preview 06-05Google83.50-33
Gemini 2.5 Pro Preview 05-06Google83.50-33
Claude Sonnet 4.5Anthropic82.40-33
Claude Opus 4Anthropic82.10-35
o4 MiniOpenAI81.40-36
DeepSeek V3.2DeepSeek81.30-37
Gemma 4 31BGoogle80.50-36
Gemma 4 31B (free)Google80.50-36
Muse Spark 1.1meta80.20-40
Gemini 3.6 FlashGoogle80.00-38
Qwen3.5 397B A17BAlibaba79.40-39
R1 0528DeepSeek79.40-39
GPT-5.4 NanoOpenAI79.30-39
GPT-5.4 MiniOpenAI79.30-40
Gemini 3.1 Flash Lite PreviewGoogle79.30-41
Gemini 2.5 Flash LiteGoogle79.10-41
Gemini 2.5 FlashGoogle79.10-42
Gemini 3.5 FlashGoogle79.00-43
GLM 5.2Zhipu AI78.50-44
GLM 5.1Zhipu AI78.00-32
GLM 5 TurboZhipu AI78.00-8
MiniMax M2.5MiniMax78.00-48
GLM 5Zhipu AI78.00-48
Qwen3.5-122B-A10BAlibaba77.70-46
Gemma 2 27BGoogle77.40-46
Qwen3.5-27BAlibaba77.00-45
Gemini 3.5 Flash LiteGoogle76.90-45
MiMo-V2.5-ProXiaomi76.20-47
Qwen3.5-35B-A3BAlibaba76.00-47
Qwen3.7 PlusAlibaba75.90-50
Kimi K2.6Moonshot AI75.80-48
o3 MiniOpenAI75.30-47
GLM 4.7Zhipu AI75.10-33
GLM 4.6Zhipu AI75.10-21
GLM 4.5Zhipu AI75.10-49
Qwen3.7 MaxAlibaba74.80+71
Qwen3.5 Plus 2026-04-20Alibaba74.80+83
Qwen3.6 Max PreviewAlibaba74.80-51
Claude Sonnet 4Anthropic74.40-50
o1OpenAI74.40-50
MiniMax M3MiniMax74.30-51
R1DeepSeek74.20-52
Qwen3.6 PlusAlibaba74.10-55
Inklingthinkingmachines74.00-54
Hy3Tencent74.00-61
o1-proOpenAI73.60-55
MiMo-V2.5Xiaomi73.00-56
Gemma 4 26B A4B Google73.00-56
Gemma 4 26B A4B (free)Google73.00-56
GLM 5V TurboZhipu AI72.3+1-56
o4 Mini HighOpenAI72.1+1-56
MiniMax M2.1MiniMax72.0+1-44
MiniMax M2MiniMax72.0+1-57
DeepSeek V3.2 ExpDeepSeek71.8+1-48
DeepSeek V3.1DeepSeek71.8+1-45
DeepSeek V3 0324DeepSeek71.8+1-60
Mistral Medium 3.5Mistral AI71.7+1-60
GPT-4o (2024-08-06)OpenAI71.20-61
GPT-4oOpenAI71.20-61
GPT-4o (2024-05-13)OpenAI71.20-61
MiniMax M1MiniMax70.80-60
GLM 4.5 AirZhipu AI70.70-60
Claude Haiku 4.5Anthropic69.50-56
DeepSeek V3DeepSeek69.50-56
Qwen3 VL 235B A22B ThinkingAlibaba69.40-47
Qwen3 VL 235B A22B InstructAlibaba69.40-57
DeepSeek V3.1 TerminusDeepSeek69.30-57
GPT-4o-miniOpenAI69.30-57
MiniMax M2-herMiniMax69.10-56
Qwen3.5-FlashAlibaba68.60-56
Hy3 previewTencent68.20-59
Qwen3.5 Plus 2026-02-15Alibaba68.20+55
Qwen3 Max ThinkingAlibaba68.20-58
Llama 4 MaverickMeta67.90-58
GPT-4.1OpenAI67.70-58
Qwen3 MaxAlibaba67.40-58
Mistral Large 3 2512Mistral AI67.00-58
Qwen3 Next 80B A3B ThinkingAlibaba66.90-41
Qwen3 Next 80B A3B InstructAlibaba66.90-59
Llama 3.3 70B InstructMeta66.80-59
GPT-4 TurboOpenAI66.70-59
Qwen3.5-9BAlibaba66.50-59
Step 3.5 FlashStepFun66.10-59
Mistral Large 2407Mistral AI65.90-33
Mistral LargeMistral AI65.90-60
Composer 2Cursor65.70-60
Composer 2 FastCursor65.70-60
GLM 4.6VZhipu AI65.50-60
Qwen3 235B A22B Thinking 2507Alibaba65.50-60
Llama 3.1 70B InstructMeta65.30-60
GPT-4 Turbo PreviewOpenAI64.80-45
GPT-4OpenAI64.80-61
Qwen3 235B A22B Instruct 2507Alibaba64.70-61
Qwen3 30B A3B Thinking 2507Alibaba64.10-59
Qwen3 30B A3B Instruct 2507Alibaba64.10+65
Qwen3 30B A3BAlibaba64.10-63
o3 Mini HighOpenAI64.10-63
Trinity Large Thinkingarcee-ai63.90-62
GPT-5 MiniOpenAI63.90-61
GLM 4.7 FlashZhipu AI63.70-61
Mixtral 8x22B InstructMistral AI63.40-61
GLM 4.5VZhipu AI62.50-61
Mercury 2Inception61.00-63
Qwen3 8BAlibaba61.00-63
Nova 2 LiteAmazon60.50-63
Phi 4Microsoft60.20-63
Kimi K2.5Moonshot AI59.10-62
gpt-oss-20bOpenAI57.40-64
gpt-oss-20b (free)OpenAI57.40-64
GPT-4o-mini (2024-07-18)OpenAI56.50-64
Granite 4.1 8BIBM55.30-63
Olmo 3 32B ThinkAllen AI54.90-63
Llama 4 ScoutMeta54.70-63
Kimi K2.7 CodeMoonshot AI54.10-63
Qwen3 235B A22BAlibaba54.00-64
GPT-4.1 MiniOpenAI53.60-64
Kimi K2 ThinkingMoonshot AI53.30-64
Kimi K2 0905Moonshot AI52.40-63
Kimi K2 0711Moonshot AI51.40-63
Claude 3 HaikuAnthropic51.30-63
Command ACohere50.80-63
Command R (08-2024)Cohere48.7+1-62
Command R+ (08-2024)Cohere48.70-63
GPT-5 NanoOpenAI46.90-63
Llama 3.1 8B InstructMeta44.50-63
GPT-4.1 NanoOpenAI42.10-63
R1 Distill Llama 70BDeepSeek40.70-63
gpt-oss-120bOpenAI40.40-63
Laguna S 2.1poolside40.00-63
Laguna S 2.1 (free)poolside40.00-63
LongCat 2.0Meituan40.00-63
Kimi K3Moonshot AI40.00-63
KAT-Coder-Air V2.5Kuaishou40.00-63
KAT-Coder-Pro V2.5Kuaishou40.00-63
Grok Latest~x-ai40.00-63
Aion-3.0-Miniaion-labs40.00-63
Aion-3.0aion-labs40.00-63
Laguna XS 2.1poolside40.00-63
Laguna XS 2.1 (free)poolside40.00-63
Fugu Ultrasakana40.00-63
North Mini Code (free)Cohere40.00-63
Claude Fable Latest~anthropic40.00-63
Nemotron 3.5 Content Safety (free)NVIDIA40.00-63
Nemotron 3 UltraNVIDIA40.00-63
Nemotron 3 Ultra (free)NVIDIA40.00-64
Step 3.7 FlashStepFun40.00-64
Perceptron Mk1perceptron40.00-63
Ring-2.6-1Tinclusionai40.00-63
GPT Chat LatestOpenAI40.00-63
Nemotron 3 Nano Omni (free)NVIDIA40.00-63
Anthropic Claude Haiku Latest~anthropic40.00-63
OpenAI GPT Mini Latest~openai40.00-63
Google Gemini Pro Latest~google40.00-63
MoonshotAI Kimi Latest~moonshotai40.00-63
Google Gemini Flash Latest~google40.00-63
Anthropic Claude Sonnet Latest~anthropic40.00-63
OpenAI GPT Latest~openai40.00-63
Qwen3.6 FlashAlibaba40.00-62
Qwen3.6 35B A3BAlibaba40.00-62
Qwen3.6 27BAlibaba40.00-62
Ling-2.6-1Tinclusionai40.00-63
Ling-2.6-flashinclusionai40.00-63
Claude Opus Latest~anthropic40.00-63
Lyria 3 Pro PreviewGoogle40.00-63
Lyria 3 Clip PreviewGoogle40.00-63
KAT-Coder-Pro V2Kuaishou40.00-63
Reka Edgerekaai40.00-63
Mistral Small 4Mistral AI40.00-63
Nemotron 3 SuperNVIDIA40.00-63
Nemotron 3 Super (free)NVIDIA40.00-63
Seed-2.0-LiteByteDance40.00-63
Seed-2.0-MiniByteDance40.00-63
Aion-2.0aion-labs40.00-63
Qwen3 Coder NextAlibaba40.00-62
Solar Pro 3Upstage40.00-62
Palmyra X5Writer40.00-62
GPT AudioOpenAI40.00-62
GPT Audio MiniOpenAI40.00-62
Seed 1.6 FlashByteDance40.00-65
Seed 1.6ByteDance40.00-65
Nemotron 3 Nano 30B A3BNVIDIA40.00-65
Nemotron 3 Nano 30B A3B (free)NVIDIA40.00-65

Which Models Are Noisy vs. Consistent?

Some models have naturally stable scores - even a small rank change for these models is meaningful. Others have volatile scores that bounce around - they need a bigger shift before you should care. CV% (coefficient of variation) tells you how volatile each model is. Higher = noisier.

Noisiest Models(highest CV% - widest significance thresholds)

ModelScoreCV%
Gemma 2 27BGoogle77.448.2%
Claude Opus 5 (Fast)Anthropic95.144.5%
Claude Opus 5Anthropic95.144.5%
Phi 4Microsoft60.243.5%
R1DeepSeek74.239.0%
Command ACohere50.836.2%
GPT-4OpenAI64.836.0%
Mixtral 8x22B InstructMistral AI63.435.5%
Claude Sonnet 5Anthropic85.235.4%
Llama 3.1 70B InstructMeta65.335.4%
Llama 3.3 70B InstructMeta66.835.1%
Mistral LargeMistral AI65.935.0%
MiniMax M2-herMiniMax69.134.9%
o3 MiniOpenAI75.333.9%
DeepSeek V3DeepSeek69.533.5%
GPT-4 Turbo PreviewOpenAI64.832.0%
Mistral Large 2407Mistral AI65.931.2%
GPT-4o (2024-05-13)OpenAI71.230.9%
GLM 5V TurboZhipu AI72.330.6%
GPT-4oOpenAI71.230.3%

Most Consistent Models(lowest CV% - tightest significance thresholds)

ModelScoreCV%
Laguna S 2.1poolside40.00.0%
Laguna S 2.1 (free)poolside40.00.0%
LongCat 2.0Meituan40.00.0%
KAT-Coder-Air V2.5Kuaishou40.00.0%
KAT-Coder-Pro V2.5Kuaishou40.00.0%
Grok Latest~x-ai40.00.0%
Aion-3.0-Miniaion-labs40.00.0%
Aion-3.0aion-labs40.00.0%
Laguna XS 2.1poolside40.00.0%
Laguna XS 2.1 (free)poolside40.00.0%
Fugu Ultrasakana40.00.0%
North Mini Code (free)Cohere40.00.0%
Claude Fable Latest~anthropic40.00.0%
Nemotron 3.5 Content Safety (free)NVIDIA40.00.0%
Nemotron 3 UltraNVIDIA40.00.0%
Nemotron 3 Ultra (free)NVIDIA40.00.0%
Step 3.7 FlashStepFun40.00.0%
Perceptron Mk1perceptron40.00.0%
Ring-2.6-1Tinclusionai40.00.0%
GPT Chat LatestOpenAI40.00.0%

How Significance Is Calculated

Understanding the statistical methodology behind our significance analysis helps you distinguish real performance shifts from random fluctuations.

Statistical Significance

We use z-scores with a 95% confidence threshold (|z| > 1.96). A z-score measures how many standard deviations a model's current score is from its historical baseline. Only changes exceeding 1.96 standard deviations are flagged as statistically significant.

Baseline Score

The baseline is computed as the arithmetic mean of each model's 14-day sparkline data. This rolling average smooths out daily fluctuations and provides a stable reference point for detecting meaningful deviations.

Confidence Intervals

Each model's 95% confidence interval is calculated as baseline ± 1.96 × standard deviation. Scores falling outside this range indicate a statistically meaningful change. The "Confidence" column shows the ± threshold value.

Multi-Timeframe Analysis

Daily (24h) and weekly (7d) rank changes are analyzed separately. Daily significance requires a rank shift of more than 3 positions; weekly requires more than 5. Models significant on both timeframes represent the strongest, most reliable signals.

Noise vs. Signal

The coefficient of variation (CV%) measures relative volatility. High-CV models have naturally noisy scores and require larger absolute changes to be significant. Low-CV models are more predictable, so even small deviations may represent real shifts.

Related

Frequently Asked Questions

Statistical significance indicates whether a model's rank change represents a real performance shift or is just random noise. We use z-scores with a 95% confidence threshold (|z| > 1.96), meaning a change is only flagged as significant if there is less than a 5% chance it occurred by random variation.

A z-score measures how many standard deviations a model's current score deviates from its historical baseline. It is calculated as (current score - baseline mean) / standard deviation. Values above +1.96 indicate significant improvement, while values below -1.96 indicate significant decline.

The CV% measures a model's relative score volatility. A high CV% means the model's performance fluctuates a lot, requiring larger changes to be statistically significant. A low CV% means the model is very consistent, so even small deviations may represent meaningful shifts. This helps distinguish inherently noisy models from truly changing ones.

AI Model Change Significance - Statistical Analysis | LM Market Cap