Comparing AI and conventional physics-based models
The models scored on this page. Every number below is the same score the leaderboard keeps: a model's daily high against what a station recorded, at one, three and seven days out, over 2026-08-08 to 2026-09-06.
Two organizations run an AI model beside their physics one, from the same starting conditions, so within a panel only the model differs. That is the comparison this page exists to make; the rows below put both in the full field.
Single-run rows are scored against the leaderboard's station highs; ensemble rows are means scored beside means and never beside the single runs. Typical miss is the mean absolute error of the daily high, pooled across every scored city and weighted by days.
The roster
What each column is, before any number. An ensemble is scored as a mean against the other means, never as a model line.
| Model | Produced by | Type | Ensemble |
|---|---|---|---|
| AIFS Single | ECMWF (Europe) | AI | no |
| AIGFS | NOAA (US) | AI | no |
| GEM | ECCC (Canada) | physics | no |
| GFS | NOAA (US) | physics | no |
| ICON | DWD (Germany) | physics | no |
| IFS | ECMWF (Europe) | physics | no |
| Unified Model | Met Office (UK) | physics | no |
| GEFS | NOAA (US) | physics | yes — 31 members |
| AIGEFS | NOAA (US) | AI | yes — 31 members |
| ECMWF ENS | ECMWF (Europe) | physics | yes — 51 members |
| AIFS ENS | ECMWF (Europe) | AI | yes — 51 members |
| WeatherNext 2 | Google DeepMind | AI | yes — 64 members |
How we approach model scoring
Verifying a weather model is typically done by comparing the forecasts produced by that model with one of several things: another weather model, a historical re-analysis, or actual surface observations. All these methods have limitations. However, we believe true model skill at the surface can best be determined by using the actual observations that occurred, instead of a computed matrix, because computed data is usually less accurate than corresponding observations, when those are available. So this page scores against stations that produce observations: 30 of them, spread across the US and Europe.
Typical miss, by lead
Mean absolute error of the daily high, pooled across every scored city and weighted by days. One scale for all three rows, so the pack visibly widens with lead. Lower is better.
AI model physics model best at that lead. Dashed rules are the AI (dark) and physics (grey) medians. Hover a dot for the model, its miss and how many model-days it rests on.
At 1 day: ICON 1.14°C, UKMO 1.25°C, GEM 1.38°C, ECMWF 1.4°C, AIGFS 1.41°C (AI), AIFS 1.56°C (AI), GFS 1.58°C. Ai median 1.48°c, physics median 1.38°c over 900–900 model-days.
At 3 days: ICON 1.4°C, UKMO 1.59°C, AIGFS 1.64°C (AI), ECMWF 1.66°C, GEM 1.69°C, AIFS 1.74°C (AI), GFS 1.81°C. Ai median 1.69°c, physics median 1.66°c over 900–900 model-days.
At 7 days: GEM 2.18°C, ECMWF 2.22°C, AIFS 2.23°C (AI), AIGFS 2.36°C (AI), GFS 2.37°C. Ai median 2.29°c, physics median 2.22°c over 900–900 model-days.
Who gets crowned, against chance
A city crowns a model only when it beats the runner-up by more than 0.22°C — the same floor the leaderboard uses. "Chance" is the share of scored models that are AI, so an AI win rate above it is evidence and one at it is not.
| Lead | AI crowned | Physics crowned | Too close | Cities | Chance (AI) |
|---|---|---|---|---|---|
| 1 day out | 2 | 3 | 25 | 30 | 29% |
| 3 days out | 4 | 4 | 22 | 30 | 29% |
| 7 days out | 5 | 7 | 18 | 30 | 40% |
How far out, and what each model publishes
Reach is the furthest day any city's newest archived forecast carries a high for the model. The AI models are thinner than their headlines: neither publishes a rain chance, a gust or an instability figure, so on those the physics models are not being beaten, they are alone. gusts and CAPE: measured 2026-08-15 (probe_model_cape.py, 6 cities, 16 leads).
| Model | Kind | Reach | Rain chance | Gusts | CAPE |
|---|---|---|---|---|---|
| AIGFS | AI | 16 d | yes | no | no |
| GFS | physics | 16 d | yes | yes | yes |
| AIFS | AI | 15 d | yes | no | no |
| ECMWF | physics | 15 d | yes | yes | yes |
| GEM | physics | 10 d | yes | yes | yes |
| ICON | physics | 7 d | yes | yes | yes |
| UKMO | physics | 7 d | yes | yes | yes |
How often each model changes its mind
Between one sweep and the next, how often a model's forecast high for the same day moved by more than 0.56°C. Pooled over 894 consecutive six-hour steps across the archived cities since 2026-08-24. Steady is not the same as right — read this beside the misses above, never alone.
AI model physics model. Dashed rules are the AI (dark) and physics (grey) medians. Hover a dot for the model's rate and its typical revision.
The full table: typical revision, how far it moved when it did, and the unchanged share
| Model | Kind | Moved > 0.56°C | Typical revision | When it moved | Unchanged | Steps |
|---|---|---|---|---|---|---|
| AIFS | AI | 18% | 0.22°C | 0.89°C | 7% | 894 |
| AIGFS | AI | 24% | 0.28°C | 0.83°C | 8% | 894 |
| ICON | physics | 39% | 0.44°C | 1.0°C | 5% | 894 |
| GFS | physics | 40% | 0.39°C | 1.06°C | 7% | 894 |
| ECMWF | physics | 42% | 0.5°C | 1.06°C | 3% | 894 |
| UKMO | physics | 45% | 0.5°C | 1.11°C | 4% | 894 |
| GEM | physics | 46% | 0.5°C | 1.39°C | 5% | 894 |
| Model | Kind | Moved > 0.56°C | Typical revision | When it moved | Unchanged | Steps |
|---|---|---|---|---|---|---|
| AIFS | AI | 30% | 0.33°C | 0.94°C | 4% | 894 |
| AIGFS | AI | 37% | 0.39°C | 1.0°C | 6% | 894 |
| ECMWF | physics | 48% | 0.56°C | 1.17°C | 3% | 894 |
| GEM | physics | 53% | 0.61°C | 1.33°C | 8% | 894 |
| ICON | physics | 54% | 0.64°C | 1.22°C | 5% | 894 |
| UKMO | physics | 56% | 0.72°C | 1.33°C | 15% | 894 |
| GFS | physics | 58% | 0.67°C | 1.22°C | 6% | 894 |
| Model | Kind | Moved > 0.56°C | Typical revision | When it moved | Unchanged | Steps |
|---|---|---|---|---|---|---|
| GEM | physics | 62% | 1.06°C | 1.97°C | 15% | 894 |
| AIFS | AI | 63% | 0.89°C | 1.5°C | 2% | 894 |
| ECMWF | physics | 65% | 1.11°C | 2.11°C | 14% | 894 |
| AIGFS | AI | 70% | 1.19°C | 1.92°C | 3% | 894 |
| GFS | physics | 78% | 1.53°C | 2.06°C | 2% | 894 |
"Unchanged" is the share of steps where the forecast did not move at all — high for a model whose newest run had not yet arrived between two sweeps, so a low move rate with a low unchanged share is a model revising in small steps rather than one standing still. Steps where a column changed which model it meant are left out.
The ensembles, mean against mean
NOAA and ECMWF each run an AI ensemble beside their physics one — NOAA's AIGEFS beside GEFS, ECMWF's AIFS ENS beside its ENS — and Google's WeatherNext 2 is an ensemble with no single run at all. What we archive for each is the members' mean and spread. A mean of many runs flattens the afternoon peak, so none of them is scored beside the single-run models above; they are scored beside each other, where the flattening cancels and, within one organization, only the model differs. None enters "the models" anywhere else on this site.
Bias of each ensemble mean's daily high against the station, by lead: warm is a mean that runs too warm, cool too cool, on one scale of ±2.78°C. The typical miss is in the table beneath.
| Ensemble | Organization | Type | 1 d miss | bias | 3 d miss | bias | 7 d miss | bias | Days |
|---|---|---|---|---|---|---|---|---|---|
| GEFS | NOAA | physics | 1.81°C | +0.58°C | 1.87°C | +0.63°C | 2.14°C | +0.42°C | 617 |
| AIGEFS | NOAA | AI | 2.18°C | -1.73°C | 2.5°C | -2.01°C | 3.01°C | -2.33°C | 617 |
| ECMWF ENS | ECMWF | physics | 1.37°C | -0.68°C | 1.56°C | -0.59°C | 1.96°C | -0.92°C | 617 |
| AIFS ENS | ECMWF | AI | 1.49°C | -1.02°C | 1.59°C | -0.94°C | 1.98°C | -1.07°C | 617 |
| WeatherNext 2 | AI | 2.16°C | -1.88°C | 2.35°C | -1.99°C | 2.7°C | -2.06°C | 637 |
GEFS since 2026-07-30; AIGEFS since 2026-07-30; ECMWF ENS since 2026-07-30; AIFS ENS since 2026-07-30; WeatherNext 2 since 2026-07-29. A forecast is scored once its day has passed and the station has reported.
Two things to hold in mind before reading a bias here. First, these are highs of an averaged series, and averaging members that peak at different hours lowers the peak a little: for GEFS the mean's high sits about half a degree below the median of its members' highs, more at longer leads. That is the size of the effect, so a lean of several degrees is the ensemble's own. Second, this is one month of one season. A lean this consistent will not average out with more days, but its size may move with the weather. Read the rows against each other, and in particular compare each AI ensemble with the physics ensemble from the same organization, since those share their starting conditions and differ only in the model. Never compare these rows with the single-run models above.
How wide each ensemble is right now
Each ensemble's member spread (one standard deviation) at the valid date's daily high, averaged across cities, at the same lead from today's newest run.
| Ensemble | Type | 1 d spread | 3 d spread | 7 d spread | Cities |
|---|---|---|---|---|---|
| GEFS | physics | 1.19°C | 1.57°C | 2.88°C | 30 |
| AIGEFS | AI | 1.44°C | 1.71°C | 3.13°C | 30 |
| ECMWF ENS | physics | 1.03°C | 1.56°C | 2.66°C | 30 |
| AIFS ENS | AI | 1.09°C | 1.69°C | 3.01°C | 30 |
| WeatherNext 2 | AI | 1.24°C | 1.67°C | 2.99°C | 30 |
Method
Scores are the leaderboard's: each model's daily high, taken from the run it had issued N days before, against the station's observed high, pooled across cities and weighted by scored days. "AI median" and "physics median" are the medians of the pooled per-model misses in each group. A model's bias has no sign until you name what it was scored against; every sign on this page is against a thermometer. Sample sizes are printed because station-scored days began in late August 2026 and a calendar span is not evidence. The National Blend of Models and our own blend are consensus products and are excluded from every count.