Why your forecast should be twenty-one forecasts
An ensemble forecast does not give you a number, it gives you a range — and the range is what tells you whether you can plan your week around it. How it works, and how to read one.
You look up Friday's forecast. It says 22 degrees. But the question you actually have is not "what will the temperature be", it is "can I plan the day around this". A single number does not answer that.
There is a deep reason for it, and it has nothing to do with how good the models are.
The atmosphere forgets where it was going
In 1961 Edward Lorenz restarted a weather simulation on his computer. To save time he began from the middle, retyping the numbers the machine had printed — rounded to three decimal places instead of six. The run, which should have reproduced the previous one, drifted away over a few simulated days until the two had nothing in common.
The starting difference was one thousandth of a degree. Not a mistake: a difference smaller than the best thermometer in the world can measure.
This is the defining property of the atmosphere: it amplifies small differences. Two initial states no instrument could tell apart produce, a week later, two different kinds of weather. It is not a flaw in the model, it is a property of the fluid. No larger computer makes it go away.
And we never know the exact state of the atmosphere. We have measurements of it — ground stations, radiosondes, airliners, buoys, and above all satellites. Millions of observations a day, and even so, between any two of them there is a gap a calculation has to fill. The starting state is always an estimate.
Running a model once from one estimate answers the question "what happens if my estimate is exactly right" — when we know it is not.
Run the model twenty-one times
The idea behind ensemble forecasting fits in a sentence: since we do not know the starting state exactly, run several forecasts from several plausible starting states, and look at what they agree on.
In practice a forecasting centre prepares:
- a control member, started from the best estimate available;
- twenty or fifty perturbed members, each started from a slightly different state, the differences chosen to resemble the real uncertainty in the observations rather than picked at random.
Many ensembles perturb the model itself as well, not only its starting point. The physics of clouds, turbulence and convection is represented by approximations, and those approximations are uncertain too. So they are varied from one member to the next.
What comes out is not a forecast. It is a population of forecasts.
The twenty-one traces start almost on top of each other. On day one they agree to within half a degree: at that range the forecast is largely determined by the current state, which we know well. Then they part. By day six they cover ten degrees.
That spread is the answer to the question you were really asking. It does not say what the weather will be. It says how open the question still is.
Count them, rather than look at them
The plume is a fine picture, but it is not easy to read. What is easy to read is the count.
Take the limit that matters to you — freezing, gusts over 60 km/h, 5 mm of rain — and count how many members cross it.
A third of the members above the limit is "it could happen, plan for it". Twenty of twenty-one is "near enough certain". Zero of twenty-one is "I can plan".
This is where the percentages in weather apps come from. A "30% chance of rain" is not a meteorologist's opinion: it is, near enough, the proportion of ensemble members that rain on you.
One caveat worth carrying: that proportion is not exactly a probability. Ensembles often understate their own uncertainty, because the perturbations applied to them do not cover every source of error. Centres correct for this by calibrating against years of archives. Take the order of magnitude rather than the decimal: three members out of twenty-one is not nothing, and it is not a certainty.
The mean is a trap
It is tempting to average the members and show only that. Plenty of interfaces do, and it is not unreasonable: on scores like root-mean-square error, the ensemble mean regularly beats any single member, including the control.
But the ensemble mean is not a possible weather. If half the members push a front through at noon and the other half at midnight, the mean drizzles all day — which no member forecasts and which will not happen.
It also smooths away the extremes, exactly where they matter. A 110 km/h gust forecast by three members out of twenty-one vanishes completely from a mean. That is precisely the information you need if you are putting up a crane.
Use the mean to get a sense of things. Use the spread to decide.
What this changes in practice
The real contribution of an ensemble is not better forecasts. It is telling you when not to trust the forecast.
There are weeks when all twenty-one members tell the same story out to Thursday. In those weeks, commit: the four-day forecast is worth what a two-day forecast would be worth otherwise.
And there are weeks — often when a low is undecided about its track — where the members separate by tomorrow. The mean forecast looks exactly as it did the week before. It is the spread, and only the spread, that tells you this one is worthless.
Three habits are enough:
- Look at the spread before you look at the value. Tight: you can plan. Wide: note the range and check again tomorrow.
- Count the members that cross your limit, rather than comparing one number to a threshold. That count is the only one that answers your question.
- Distrust a forecast that never changes. A seven-day forecast that reads identically three days running is more often a frozen display than a decided atmosphere.
The ensembles you will meet
| Ensemble | Centre | Members | Reach |
|---|---|---|---|
| ENS | ECMWF | 50 + control | 15 days |
| GEFS | NOAA | 30 + control | 16 days |
| GEPS | Environment Canada | 20 + control | 16 days |
| PEARP | Météo-France | 34 + control | 108 hours |
| AROME-EPS | Météo-France | 16 | 51 hours |
The last two are worth a word: they are fine-mesh ensembles, built for the uncertainty that matters at short range — will the storm cross this parish or the next one. A global ensemble on an 18 km mesh cannot answer that, because it cannot see the storm.
One last thing
An ensemble forecast does not tell you what will happen. It tells you how much the atmosphere still knows about itself at that range — and that is honest information, which the single number in your weather app hides from you in the name of simplicity.
A single number is never wrong. It is just rarely right. A spread tells you how much to believe it.