Skip to content

Change Significance Tracker

Not all rank changes are meaningful. Some are random noise. This page uses statistical analysis to tell you which model score movements are real trends vs. normal fluctuation, so you know which changes to pay attention to.

Models Analyzed

278

Significant Changes

17

Noise (Not Significant)

261

Both Timeframes

48

What This Means

  • 17 of 278 models have score changes that are statistically significant - these are real performance shifts, not random noise.
  • 261 models show score changes within their normal variation range - don't read too much into small rank shifts for these models.
  • 48 models show significant changes on both daily and weekly timeframes - these are the strongest, most reliable signals of real performance change.
  • How to use this: When a model's rank changes, check here first. If it's not flagged as significant, the change is likely temporary noise. If it is significant, the model is genuinely improving or declining.

Top Significant Changes (by |Z-Score|)

LMMarketCap.com

Real Performance Shifts

17 models whose recent scores deviate enough from their historical average to be considered a real change (not noise). Sorted by how extreme the change is. Z-Score measures how unusual the change is - values beyond ±1.96 mean there's a 95% chance the change is real.

ModelCurrentZ-ScoreDirection
Qwen3 30B A3B Instruct 2507Alibaba64.15.04Improvement
Qwen3.5 Plus 2026-04-20Alibaba74.84.58Improvement
GLM 5.1Zhipu AI78.04.29Improvement
Qwen3.7 MaxAlibaba74.84.24Improvement
Claude Sonnet 5Anthropic85.23.46Improvement
gpt-oss-120bOpenAI61.63.02Improvement
Claude Opus 5Anthropic95.13.00Improvement
Kimi K2.5Moonshot AI47.5-2.99Decline
Claude Fable 5Anthropic95.9-2.68Decline
Grok 4.3xAI88.82.65Improvement
Muse Spark 1.2meta80.92.65Improvement
Grok 4.5xAI88.82.64Improvement
Grok 4.6xAI88.82.45Improvement
Claude Fable 5 (batch)Anthropic95.9-2.45Decline
GPT-4o (2024-11-20)OpenAI71.22.41Improvement
GPT-5 MiniOpenAI53.8-2.10Decline
GPT-5 NanoOpenAI61.92.01Improvement

Short-Term vs. Sustained Changes

A model changing rank in 24 hours could be a blip. But if it's also moving over 7 days, that's a real trend. Models flagged on both timeframes are the most important to watch - they represent confirmed, sustained performance shifts.

Significant on Both Timeframes(strongest signals)

ModelScore24h Change7d Change
Gemini 3.6 FlashGoogle40.0-9-13
Kimi K3Moonshot AI40.0-9-13
Gemini 3.6 Flash (batch)Google40.0-9-13
Schematron V2 Turboinference-net40.0-12-21
Schematron V2 Smallinference-net40.0-12-21
GPT Astra Latest~openai40.0-12-21
GPT Sol Latest~openai40.0-12-21
GPT Terra Latest~openai40.0-12-21
GPT Luna Latest~openai40.0-12-21
Fugu Ultra v2sakana40.0-12-21
Fugu Maxsakana40.0-12-21
Ling 3.0 Flash VLinclusionai40.0-12-21
Ling 3.0 Flash VL (free)inclusionai40.0-12-21
DeepSeek V4.1 FlashDeepSeek40.0-12-21
Ling 3.0 Flash Sante (free)inclusionai40.0-13-22
Muse Spark 1.3 Contributormeta40.0-13-21
Hy4 previewTencent40.0-13-20
Ling 3.0 Flash Fininclusionai40.0-13-20
Ling 3.0 Flash Fin (free)inclusionai40.0-13-20
GLM Flash Latest~z-ai40.0-13-20
Qwen3.8 FlashAlibaba40.0-13-20
Muse Spark 1.2 Contributormeta40.0-13-20
DeepSeek V4 Flash Vision ExpDeepSeek40.0-13-20
Hy-MT2-1.8BTencent40.0-12-19
Hy-MT2-30B-A3BTencent40.0-12-19
GLM Latest~z-ai40.0-12-19
Hy-MT2-7BTencent40.0-12-19
Dots3-Note Preview (free)dots-studio40.0-12-19
Gemini 3.7 FlashGoogle40.0-12-19
Gemini 3.7 Flash (batch)Google40.0-12-19
Seed 2.1 TurboByteDance40.0-12-19
Qwen3.8 2.4T A95BAlibaba40.0-12-19
Seed-2.0-CodeByteDance40.0-12-18
DeepSeek V4 Pro 0813DeepSeek40.0-12-18
LFM2.5-2.6B (free)Liquid AI40.0-11-17
Nemotron 3.5 LightningNVIDIA40.0-11-17
Nemotron 3.5 Lightning (free)NVIDIA40.0-11-17
Sakana Namazusakana40.0-11-17
Solar Pro 4Upstage40.0-11-17
Muse Glimmer 30Bmeta40.0-11-17
DeepSeek V4 Flash Latest~deepseek40.0-10-15
DeepSeek V4 Flash 0731DeepSeek40.0-10-15
Qwen3.7 FlashAlibaba40.0-9-14
Claude Opus 5 (batch)Anthropic40.0-9-13
Ling 3.0 Flashinclusionai40.0-9-13
Laguna S 2.1poolside40.0-9-13
Laguna S 2.1 (free)poolside40.0-9-13
LongCat 2.0Meituan40.0-9-13

Daily Only(may be noise)

No models with daily-only significance.

Weekly Only(building trend)

ModelScore24h Change7d Change
Claude Opus 5Anthropic95.1-1+272
GPT-5.5 ProOpenAI92.7-1+8
GPT-5.3-CodexOpenAI90.5-1+18
GPT-5.2 ChatOpenAI90.5-1+80
Grok 4.5xAI88.8-1+56
Grok 4.3xAI88.8-1+30
Grok 4.20xAI88.8-1-8
GPT-5 ProOpenAI88.7-1-10
GPT-5 Pro (batch)OpenAI88.7-1-10
GPT-5OpenAI88.7-1-10
GPT-5 (batch)OpenAI88.7-1-10
Gemini 3 Flash PreviewGoogle88.4-1-10
Gemini 3 Flash Preview (batch)Google88.4-1-10
Grok 4.20 Multi-AgentxAI87.9-1-9
GPT-5.1 (batch)OpenAI87.8-1-7
GPT-5.1-Codex-MiniOpenAI87.8-1-6
o3 ProOpenAI86.7-1-6
o3OpenAI86.7-1-6
o3 (batch)OpenAI86.7-1-6
Claude Sonnet 5Anthropic85.2-1+231
Claude Sonnet 4.6Anthropic85.2-1-7
Claude Sonnet 4.6 (batch)Anthropic85.2-1-7
Claude Opus 4.5Anthropic85.1-1-7
Claude Opus 4.5 (batch)Anthropic85.1-1-7
Gemini 2.5 ProGoogle83.5-1-7
Gemini 2.5 Pro (batch)Google83.5-1-7
Gemini 2.5 Pro Preview 06-05Google83.5-1-7
DeepSeek V3.2DeepSeek83.4-1-7
Claude Sonnet 4.5Anthropic82.4-1-7
Claude Sonnet 4.5 (batch)Anthropic82.4-1-7
GPT-6 AstraOpenAI81.8-1-6
GPT-6 Astra (batch)OpenAI81.8-1-6
GPT-6 Astra ProOpenAI81.8-1-6
GPT-6 Astra Pro (batch)OpenAI81.8-1-6
o4 MiniOpenAI81.4-1-6
o4 Mini (batch)OpenAI81.4-1-6
Gemini 3.8 FlashGoogle81.0-1-6
Gemini 3.8 Flash (batch)Google81.0-1-6
Muse Spark 1.3meta80.9-1+161
Muse Spark 1.2meta80.9-1+189
Muse Spark 1.1meta80.9-1-7
Gemma 4 31BGoogle80.5-1-7
Gemma 4 31B (free)Google80.5-1-7
Qwen3.5 397B A17BAlibaba79.4-1-7
R1 0528DeepSeek79.4-1-7
GPT-5.4 NanoOpenAI79.3-1-7
GPT-5.4 Nano (batch)OpenAI79.3-1-7
GPT-5.4 MiniOpenAI79.3-1-7
GPT-5.4 Mini (batch)OpenAI79.3-1-7
Gemini 3.1 Flash Lite PreviewGoogle79.3-1-7
Gemini 2.5 Flash LiteGoogle79.1-1-7
Gemini 2.5 Flash Lite (batch)Google79.1-1-7
Gemini 2.5 FlashGoogle79.1-1-7
Gemini 2.5 Flash (batch)Google79.1-1-7
Gemini 3.5 FlashGoogle79.0-1-7
Gemini 3.5 Flash (batch)Google79.0-1-7
GLM 5.3 (batch)Zhipu AI78.6-1-8
GLM 5.2Zhipu AI78.6-1-8
GLM 5.3 FlashZhipu AI78.00-7
GLM 5.1Zhipu AI78.00+11
GLM 5 TurboZhipu AI78.00+43
GLM 5Zhipu AI78.00-9
GLM 5.3 Flash (batch)Zhipu AI77.90-9
Qwen3.5-122B-A10BAlibaba77.70-8
Gemma 2 27BGoogle77.40-8
Qwen3.5-27BAlibaba77.00-6
GPT-5 Mini (batch)OpenAI76.80-6
GPT-5 Nano (batch)OpenAI76.80-6
Gemini 3.5 Flash LiteGoogle76.50-6
Gemini 3.5 Flash Lite (batch)Google76.50-6
MiMo-V2.5-ProXiaomi76.20-7
Qwen3.5-35B-A3BAlibaba76.00-7
GLM 5.2 (free)Zhipu AI75.70-7
Kimi K2.6Moonshot AI75.70-7
Qwen3.8 Max (0902)Alibaba75.60+120
Qwen3.7 PlusAlibaba75.60-8
o3 MiniOpenAI75.30-7
Claude Opus 4.1 (batch)Anthropic75.20-7
GLM 4.7Zhipu AI75.10+13
GLM 4.6Zhipu AI75.10+27
GLM 4.5Zhipu AI75.10-9
Qwen3.7 MaxAlibaba74.80+176
Qwen3.5 Plus 2026-04-20Alibaba74.80+187
Qwen3.6 Max PreviewAlibaba74.80-11
Hy3Tencent74.40-11
Claude Opus 4.1Anthropic74.40-11
o1OpenAI74.40-10
o3 Mini (batch)OpenAI74.30-10
Qwen3.6 PlusAlibaba74.10-10
MiniMax M3MiniMax73.90-10
Claude Sonnet 4Anthropic73.90-15
R1DeepSeek73.80-11
o1-proOpenAI73.60-10
Qwen3.8 27BAlibaba73.40-10
MiMo-V2.5Xiaomi73.00-10
Gemma 4 26B A4B Google73.00-10
Gemma 4 26B A4B (free)Google73.00-10
Inklingthinkingmachines72.90-10
Inkling (free)thinkingmachines72.90-10
GLM 5V TurboZhipu AI72.30-8
o4 Mini HighOpenAI72.10-8
GPT-4o (batch)OpenAI72.00-8
DeepSeek V3.1DeepSeek71.80+7
DeepSeek V3 0324DeepSeek71.80-10
Mistral Medium 3.5Mistral AI71.60-10
MiniMax M2.7MiniMax71.60+6
MiniMax M2.5MiniMax71.60-11
GPT-4o (2024-11-20)OpenAI71.20+60
GPT-4o (2024-08-06)OpenAI71.20-12
GPT-4oOpenAI71.20-12
GPT-4o (2024-05-13)OpenAI71.20-12
MiniMax M2MiniMax71.00-12
MiniMax M1MiniMax70.80-12
GLM 4.5 AirZhipu AI70.70-12
Llama 4 MaverickMeta70.70-12
Claude Haiku 4.5Anthropic69.90-9
Claude Haiku 4.5 (batch)Anthropic69.90-9
DeepSeek V3DeepSeek69.50-8
Qwen3 VL 235B A22B InstructAlibaba69.30-8
DeepSeek V3.1 TerminusDeepSeek69.30-8
GPT-4o-miniOpenAI69.30-8
Inkling Smallthinkingmachines68.70-8
Inkling Small (free)thinkingmachines68.70-8
Qwen3.5-FlashAlibaba68.60-8
Hy3 previewTencent68.40-8
Qwen3.5 Plus 2026-02-15Alibaba68.20+159
Qwen3 Max ThinkingAlibaba68.20-9
MiniMax M2-herMiniMax68.10-9
GPT-4.1OpenAI67.70-9
GPT-4.1 (batch)OpenAI67.70-9
Qwen3 MaxAlibaba67.40-8
Mistral Large 3 2512 (batch)Mistral AI67.00-8
Llama 3.3 70B InstructMeta66.80-8
Qwen3 Next 80B A3B ThinkingAlibaba66.70+6
Qwen3 Next 80B A3B InstructAlibaba66.70-9
GPT-4 TurboOpenAI66.70-9
Qwen3.5-9BAlibaba66.50-9
Step 3.5 FlashStepFun66.10-9
Mistral Large 2407Mistral AI65.90+23
Mistral LargeMistral AI65.90-10
Composer 2Cursor65.70-10
Composer 2 FastCursor65.70-10
GLM 4.6VZhipu AI65.60-10
Qwen3 235B A22B Thinking 2507Alibaba65.30-10
Llama 3.1 70B InstructMeta65.30-10
GPT-4OpenAI64.90-10
Qwen3 235B A22B Instruct 2507Alibaba64.70-10
Qwen3 30B A3B Thinking 2507Alibaba64.10-9
Qwen3 30B A3B Instruct 2507Alibaba64.10+170
Qwen3 30B A3BAlibaba64.10-12
o3 Mini HighOpenAI63.90-10
GLM 4.7 FlashZhipu AI63.50-10
Mixtral 8x22B InstructMistral AI63.40-10
Trinity Large Thinkingarcee-ai63.10-10
GPT-4o-mini (batch)OpenAI62.50-10
GLM 4.5VZhipu AI62.30-10
GPT-5 NanoOpenAI61.90-10
gpt-oss-120bOpenAI61.60-10
GPT-4 Turbo (batch)OpenAI61.50-9
Mercury 2.5Inception61.20-9
Qwen3 8BAlibaba61.00-9
Mercury 2Inception60.90-9
Nova 2 LiteAmazon60.40-9
Llama 4 ScoutMeta60.20-9
Phi 4Microsoft60.20-9
GPT-4.1 Mini (batch)OpenAI58.90-9
GPT-4.1 Nano (batch)OpenAI58.90-9
gpt-oss-20bOpenAI57.40-9
gpt-oss-20b (batch)OpenAI57.40+213
GPT-4o-mini (2024-07-18)OpenAI56.5-1-10
GPT-4.1 MiniOpenAI56.2-1-10
Qwen3 235B A22BAlibaba54.0-1-9
Granite 4.2 8BIBM53.8-1-9
GPT-5 MiniOpenAI53.8-1-9
Command A+Cohere53.20-7
Claude 3 HaikuAnthropic51.3-2-9
Command ACohere50.8-2-9
Command R (08-2024)Cohere48.7-1-8
Command R+ (08-2024)Cohere48.7-2-9
Kimi K2.5Moonshot AI47.5-2-9
Llama 3.1 8B InstructMeta44.5-2-9
GPT-4.1 NanoOpenAI42.1-2-9
R1 Distill Llama 70BDeepSeek40.7-2-9
Kimi K3 (batch)Moonshot AI40.00-13

Which Models Are Noisy vs. Consistent?

Some models have naturally stable scores - even a small rank change for these models is meaningful. Others have volatile scores that bounce around - they need a bigger shift before you should care. CV% (coefficient of variation) tells you how volatile each model is. Higher = noisier.

Noisiest Models(highest CV% - widest significance thresholds)

ModelScoreCV%
Gemma 2 27BGoogle77.441.4%
Phi 4Microsoft60.237.7%
Claude Opus 5Anthropic95.136.3%
Gemini 3.6 FlashGoogle40.035.3%
Muse Spark 1.3meta80.935.3%
R1DeepSeek73.834.0%
Kimi K3Moonshot AI40.032.7%
Command ACohere50.831.7%
Qwen3.8 Max (0902)Alibaba75.631.5%
GPT-4OpenAI64.931.5%
Command A+Cohere53.231.2%
Mixtral 8x22B InstructMistral AI63.431.0%
Llama 3.1 70B InstructMeta65.331.0%
Llama 3.3 70B InstructMeta66.830.8%
Mistral LargeMistral AI65.930.7%
Gemini 3.6 Flash (batch)Google40.030.6%
MiniMax M2-herMiniMax68.130.5%
Muse Spark 1.2meta80.930.0%
o3 MiniOpenAI75.329.8%
DeepSeek V3DeepSeek69.529.4%

Most Consistent Models(lowest CV% - tightest significance thresholds)

ModelScoreCV%
GPT-5.5 Pro (batch)OpenAI92.70.0%
GPT-5.5 (batch)OpenAI92.70.0%
GPT-5.2 Pro (batch)OpenAI90.50.0%
GPT-5.2 (batch)OpenAI90.50.0%
GPT-5.6 Luna Pro (batch)OpenAI89.00.0%
GPT-5.6 Luna (batch)OpenAI89.00.0%
GPT-5.6 Terra Pro (batch)OpenAI89.00.0%
GPT-5.6 Terra (batch)OpenAI89.00.0%
GPT-5.6 Sol Pro (batch)OpenAI89.00.0%
GPT-5.6 Sol (batch)OpenAI89.00.0%
GPT-5 Pro (batch)OpenAI88.70.0%
GPT-5 (batch)OpenAI88.70.0%
GPT-5.1 (batch)OpenAI87.80.0%
o3 (batch)OpenAI86.70.0%
Claude Sonnet 4.6 (batch)Anthropic85.20.0%
Gemini 2.5 Pro (batch)Google83.50.0%
Grok 4.3 (batch)xAI81.00.0%
GPT-5.4 Nano (batch)OpenAI79.30.0%
GPT-5.4 Mini (batch)OpenAI79.30.0%
Gemini 3.5 Flash (batch)Google79.00.0%

How Significance Is Calculated

Understanding the statistical methodology behind our significance analysis helps you distinguish real performance shifts from random fluctuations.

Statistical Significance

We use z-scores with a 95% confidence threshold (|z| > 1.96). A z-score measures how many standard deviations a model's current score is from its historical baseline. Only changes exceeding 1.96 standard deviations are flagged as statistically significant.

Baseline Score

The baseline is computed as the arithmetic mean of each model's 14-day sparkline data. This rolling average smooths out daily fluctuations and provides a stable reference point for detecting meaningful deviations.

Confidence Intervals

Each model's 95% confidence interval is calculated as baseline ± 1.96 × standard deviation. Scores falling outside this range indicate a statistically meaningful change. The "Confidence" column shows the ± threshold value.

Multi-Timeframe Analysis

Daily (24h) and weekly (7d) rank changes are analyzed separately. Daily significance requires a rank shift of more than 3 positions; weekly requires more than 5. Models significant on both timeframes represent the strongest, most reliable signals.

Noise vs. Signal

The coefficient of variation (CV%) measures relative volatility. High-CV models have naturally noisy scores and require larger absolute changes to be significant. Low-CV models are more predictable, so even small deviations may represent real shifts.

Related

Frequently Asked Questions

Statistical significance indicates whether a model's rank change represents a real performance shift or is just random noise. We use z-scores with a 95% confidence threshold (|z| > 1.96), meaning a change is only flagged as significant if there is less than a 5% chance it occurred by random variation.

A z-score measures how many standard deviations a model's current score deviates from its historical baseline. It is calculated as (current score - baseline mean) / standard deviation. Values above +1.96 indicate significant improvement, while values below -1.96 indicate significant decline.

The CV% measures a model's relative score volatility. A high CV% means the model's performance fluctuates a lot, requiring larger changes to be statistically significant. A low CV% means the model is very consistent, so even small deviations may represent meaningful shifts. This helps distinguish inherently noisy models from truly changing ones.

AI Model Change Significance - Statistical Analysis | LM Market Cap