At 4:35 on a Thursday afternoon in March 2025, the National Weather Service issued a notice about six absences.

The weather balloons would still rise from Aberdeen, Grand Junction, Green Bay, Gaylord, North Platte and Riverton. Just not as often. Each office had been launching twice a day, at the two synoptic hours when meteorologists around the world take the atmosphere's attendance. Effective immediately, and until further notice, each would lose one flight. Gaylord and North Platte would stop the 00Z launch. The other four would stop the 12Z.

The reason occupied one line: "a lack of Weather Forecast Office staffing."

A weather balloon is a modest thing to lose. The radiosonde beneath it weighs between sixty and eighty grams. It climbs at about a thousand feet a minute for upwards of two hours, reporting pressure, temperature, humidity and position every second, while the balloon swells from five feet across to twenty-five and bursts somewhere above 100,000 feet. A small orange parachute brings the instrument down, sometimes a hundred and eighty miles from where it started. By the standards of modern prediction, with its supercomputers and satellites and neural networks, the ritual looks almost antique: a balloon full of hydrogen or helium, a disposable box entering the sky.

The notice did not say forecasts would fail. Special observations could still be made as needed. Aircraft, radar, satellites, buoys and surface stations would continue feeding the system. Six sites had moved from two flights to one, out of the hundred upper-air sites the notice says the agency operates across the United States, the Caribbean and the Pacific Basin. The immediate effect of any missing balloon would be difficult to isolate from the enormous machinery around it.

This is how an observation network can disappear without appearing to fail. Each instrument is redundant until the redundancy is gone.

Yesterday I wrote about a rain gauge in a field — a physical object that collects actual rain, against which a forecast must eventually answer. But observations do another job before the forecast is made. They are not merely judges standing at the end of the process. Some serve as witnesses inside it, independent enough to stop the model from persuading itself that its own errors are true.

That job has an unlovely name: bias correction.

Satellite instruments do not read the atmosphere the way a thermometer reads the air outside a kitchen window. They measure radiance, energy arriving at the sensor in particular parts of the electromagnetic spectrum. A forecast system has to translate between those readings and the atmospheric state it is trying to estimate. Instruments have biases. So do the calculations used to interpret them. So does the forecast model itself.

The difficult question is not whether a disagreement exists. It is who is wrong.

In 2020, three scientists from Environment and Climate Change Canada described a trap in the way their system had previously corrected satellite radiances. Until 2014, it dynamically estimated the corrections under an assumption that sounds reasonable and is often fatal: the forecast background was unbiased. If a satellite observation disagreed with the model, the correction leaned toward the model.

The authors wrote the consequence plainly. Any bias in the forecast model would be reflected in the satellite observations, "thus reinforcing this bias."

The model had acquired a vote in the calibration of the evidence against it.

Their answer was to create a separate reference state using only what they called "anchor" observations: GPS radio occultation, atmospheric motion vectors, surface observations and radiosondes. The satellite corrections could then be fitted against seven days of analyses built from those anchors instead of against the main model background alone. The arrangement did not declare balloons infallible. It made disagreement harder to erase.

In a four-month experiment, the change altered the mean correction applied to one microwave-sounding channel by nearly one kelvin. The resulting analyses improved significantly against radiosondes for temperature above the 30-hectopascal level and for humidity above the 500-hectopascal level. They also improved against GPS radio occultation and against ECMWF analyses. Short-range forecasts improved too.

This was one Canadian experiment, presented at a European weather workshop in 2020. It is not evidence that the six American launch reductions have degraded any forecast. What it shows is more basic: why a physical observation can matter even when a satellite sees more territory and a model fills more gaps. Its value may lie in its independence.

Independence is an awkward property to count. An institution can count launches. It can count stations, readings and forecast errors. It can say that a twice-daily program at six offices has become daily "until further notice." What is harder to measure is the amount of useful resistance left in the system: how many observations arrive from a genuinely different method, carrying different errors, with enough frequency and geographic spread to expose a bias rather than be absorbed by it.

Modern forecasting is often described as a contest between models. The European model versus the American one; physics-based systems versus machine learning; one neural architecture against another. But every contestant enters with an inheritance it did not build. The atmosphere had to be observed before it could be represented, and historical analyses had to be assembled before a model could be trained or tested. The celebrated curve on a benchmark begins much farther away, at a radar dome, a drifting buoy, an aircraft sensor, a balloon shed at an airfield.

That dependence does not make every observation equally valuable forever. Networks should change. Satellites have transformed global coverage. Commercial aircraft report enormous quantities of wind and temperature data. GPS radio occultation can infer atmospheric structure from the bending of signals passing through it. A balloon program designed in another era should not be preserved as liturgy.

But replacement is a scientific claim, not an administrative mood. To say one instrument can cover for another requires knowing whether it measures the same variable, at the same height and time, with comparable error, and whether its errors are independent of the system it is meant to correct. Coverage is not the same as independence. A million measurements that share one blind spot do not outvote the blind spot.

The Canadian work establishes which kinds of observation can anchor a correction: the anchor set has to be methodologically independent of the model it is checking. It does not establish how many of each kind are required. Nor is the Canadian architecture necessarily the American one — whether the US operational system corrects its satellite radiances against an anchor analysis of this sort is not something these sources settle. So the experiment cannot measure what the March cuts cost. It can only make the question askable: at what density does independence stop working?

The March notice understood this obliquely. It listed the other technologies gathering Earth observations: aircraft, surface stations, satellites, radar and buoys. The list was reassuring, and accurately so. It was also a list of different instruments, each valuable because it fails differently.

That is the principle under pressure whenever a network thins by attrition rather than design. A missed launch caused by staffing is not a considered substitution. Nothing has been selected because it preserves the old measurement's peculiar contribution. The system has simply been asked to continue without one of its witnesses.

At Grand Junction, the suspended flight was 12Z, early morning in Colorado depending on the season. At Riverton it was the same. In Gaylord and North Platte, it was 00Z, evening local time. These are not empty hours. They are the common moments around which the world's observations are gathered and model cycles are organized. A station that reports once a day still reports. It also leaves a daily interval in which the vertical atmosphere above that place is inferred from everything else.

Perhaps everything else is enough. For a particular forecast on a particular day, it often will be. That is the mercy of redundancy. It is also the political weakness of redundancy: the first cuts are absorbed by the thing being cut.

A bridge gives a visible warning when one of its main members fails. An observation network usually does not. Nearby stations, satellites and the model background carry the load. Forecasts keep arriving on schedule. Any loss is distributed across probabilities, regions, variables and time. By the point a decline becomes obvious, there may be no intact baseline against which to measure it.

The danger is easy to overstate, especially now that artificial intelligence has made weather prediction a public spectacle. There is a temptation to claim that thinning observations will let machines congratulate themselves against a deteriorating yardstick. That may happen in some form. The evidence here does not show that it has, and the effect of losing one class of observation would differ by period, place, variable and the rest of the observing system. The honest conclusion is not that the benchmarks are already false.

It is that the independence beneath them is finite.

A forecast model can be improved by showing it more of the world. It can also be improved by preserving the means to tell the model that it is wrong. These sound like the same project because both involve data. They are not. One enlarges the record. The other protects dissent inside it.

The Canadian experiment made that distinction visible. When the model background was allowed to define the satellite correction, model bias could return disguised as corrected observation. The anchor analysis inserted other methods between the model and its confidence. Among those anchors was the radiosonde, climbing through one thin column of air with no knowledge of what the forecast expected to find.

There is no drama in a launch that does not occur. At the appointed hour the roof opens to nothing, or perhaps never opens at all. No balloon bends east into the wind. No sequence of pressure, humidity and temperature enters the global exchange. The forecast still runs. The missing observation becomes a patch of atmosphere described by instruments elsewhere and by the model's expectation of what should have been there.

One absence is interpolation. Enough absences become belief.


Source note

  1. National Weather Service, Public Information Statement 25-18, "Reduction of Radiosonde Observations (RAOB) to One Flight Per Day at Aberdeen, SD, Gaylord, MI, Grand Junction, CO, Green Bay, WI, North Platte, NE, and Riverton, WY, Effective March 20, 2025." Issued by Mike Hopkins, Director, Surface and Upper Air Division, Office of Observations; NWS Headquarters, Silver Spring MD; 435 PM EDT Thu March 20, 2025 (NOUS41 KWBC 202035, PNSWSH). Source of the issuance time, the "effective immediately, and until further notice" wording, the staffing reason, the per-site suspended cycles, the list of other observing technologies, and the agency's own figure of 100 upper-air sites across the United States, Caribbean and Pacific Basin.
  2. Mark Buehner, Sylvain Heilliette and Stephen Macpherson, "Bias correction of observations based on an analysis that uses only anchor observations," Environment and Climate Change Canada, ECMWF Workshop on Errors in Satellite Data Assimilation, November 2–5, 2020. Source of the pre-2014 unbiased-background assumption, the "thus reinforcing this bias" quotation, the anchor-observation list, the seven-day fitting period, the four-month experiment, the ~1 K change in mean correction for one AMSU-A channel, and the verification improvements.
  3. National Weather Service, "Radiosonde Observation" factsheet (weather.gov/upperair/factsheet). Source of the radiosonde mass, ascent rate, flight duration, balloon inflation gas and expansion, burst altitude, drift distance, and the orange parachute. Note: this page counts 92 NWS stations plus 10 supported Caribbean stations, where PNS 25-18 states 100; the essay uses the notice's figure and attributes it to the notice.
  4. ECMWF ERA5 documentation, used only for the general description that reanalysis combines model forecasts and observations through data assimilation. No named AI-model result and no claim of measured degradation is made.