Four days before Hurricane Melissa came ashore in Jamaica, the National Hurricane Center issued a track forecast that placed the storm's centre within 11 nautical miles of where it actually arrived.
Eleven nautical miles at 96 hours. The NHC's own report calls this "far below the typical error for a 4-day forecast," which is a restrained way of describing one of the better forecasts in the agency's history. The advisory issued at 1500 UTC on 24 October 2025, when Melissa was still a 40-knot tropical storm sitting in the central Caribbean, called for it to become a strong Category 4 hurricane and to track near or directly over western Jamaica. Subsequent forecasts gave nearly three days of warning that it would arrive as a Category 5. It was the first time the NHC had ever forecast a storm to reach that strength from an initial intensity so low. Jamaica had a tropical storm watch roughly seven days out, a hurricane watch about five days out, and a hurricane warning about three days out.
What arrived on 28 October, at 1725 UTC near New Hope in Westmoreland Parish, was a 160-knot hurricane with a central pressure of 897 millibars: the strongest landfall on record in Jamaica, and the second lowest landfall pressure ever measured in the Atlantic basin, behind the Labor Day hurricane of 1935. Storm surge between Crawford and Black River reached an estimated 7 to 11 feet.
The damage came to $12.2 billion. That is 41 percent of Jamaica's 2024 gross domestic product, and it is the figure that tells you what the three days of warning were worth.
That much is a solved problem, or close enough to solved that the remaining error is measured in single-digit nautical miles.
And then there is the rain.
The NHC report is candid about it. "Melissa's rainfall forecast proved particularly challenging." Most of the severe flooding occurred in outer rainbands well removed from the centre rather than in the inner core. Deep moisture running into steep terrain over Hispaniola and eastern Cuba concentrated into "narrow corridors of intense orographic rainfall that models struggled to resolve." Over Jamaica the early forecasts put the heaviest rain across the centre and east of the island, with local totals of 20 to 30 inches. Eastern Jamaica ended up with 6 to 14 inches. The west, where the storm actually went, took more than 24 inches at several stations and something near 32 inches at the local maximum. Forecasters extended the rainfall product from its normal three-day window to four days because the cumulative threat would not fit inside three.
At least 93 people died: 45 in Jamaica, at least 43 in Haiti, 5 elsewhere. The southwest parish of St. Elizabeth recorded the highest toll, 18, followed by Westmoreland with 15. Both are in the west, where the rain went and the early forecasts said it would not.
What the report will not do is tell you what killed them. "A breakdown of the number of direct and indirect deaths is not known." On Jamaica: "A detailed breakdown of the causes of the fatalities is not available." On Haiti: "Details on the causes of these fatalities are not known as of this writing." Three times in one section, the agency declines to make the attribution.
And then the count runs out entirely. As the floodwater receded, Jamaica confirmed an outbreak of leptospirosis, a bacterial infection carried in contaminated water and soil, and declared a health emergency. People died of it. The report says so. It also says the number is unknown, and that those deaths are not in the casualty table at all.
So the instrument stops before the consequence does. Ninety-three is not the number of people Melissa killed; it is the number of deaths the measuring system could still see when it was time to publish.
One storm, two entirely different forecasting problems, verified against each other in the same document. Where the hurricane would go was nailed to within eleven miles four days out. Where the water would fall was moved around the island between forecast cycles until the storm arrived and settled the question itself.
On 3 September 2026, Google DeepMind and Google Research announced WeatherNext 3, which they describe as the most advanced and accurate global weather model to date. It is a substantial piece of engineering and the claims around it are more specific than the usual launch material, which makes it unusually useful to think with.
The headline is resolution. WeatherNext 3 resolves surface temperature and moisture at 5 kilometres, other surface variables at 10 kilometres, and atmospheric variables such as wind speed at 25 kilometres. Google calls this "roughly five times sharper" than WeatherNext 2, which ran on a 25-kilometre grid in six-hour increments. The model now produces a new forecast every hour.
Underneath the resolution is a change in what the model learns from. Most AI weather models, including WeatherNext 2, are trained on the output of numerical weather prediction systems, the physics simulations that have run on national supercomputers for six decades. Those systems carry a six-hour data lag. WeatherNext 3 instead ingests a mosaic of live global geostationary satellite imagery and trains directly on sparse weather station observations. For precipitation it trains on NASA's IMERG satellite product and on an internal global precipitation reanalysis built from satellite radar. The stated aim is regions that have never had high-resolution forecasting because nobody would build them a regional supercomputer model: Latin America, Africa, Asia-Pacific. Those are the places where a day of warning is worth the most and where almost nobody has been paid to provide it.
The resolution figures carry a caveat the headline does not. "Five times sharper" is true for surface temperature and moisture at 5 kilometres. It is not true for wind aloft, which stayed at 25 kilometres, exactly where WeatherNext 2 left it. The gains arrived where the observations are dense, at the surface, where there are weather stations and geostationary imagery. Where the observations are thin, the resolution did not move. Even the precipitation comparison figure in the announcement shows WeatherNext 3 at 11 kilometres, not at 5.
Two things about the announcement are worth reporting carefully, because they bear on how much weight the superlative can hold.
Google's claim to be "the most advanced and accurate global weather model to date" is attributed to "independent live evaluations by Brightband." Brightband is a startup building open benchmark datasets, models and metrics for weather forecasting, and its stated purpose is to establish a common evaluation task for the field. Its people are serious: the advisory lead is Amy McGovern, who directs the NSF AI institute for trustworthy AI in weather and climate at the University of Oklahoma; the chief scientist is Ryan Keisler, previously of Descartes Labs; the head of data and weather is Daniel Rothenberg, formerly chief scientist at Tomorrow.io. The company's chief executive and co-founder, Julian Green, was previously general manager of AI moonshots at Google X, and his earlier company Jetpac was acquired by Google. Those facts are listed on Brightband's own website. They do not imply that the evaluation is wrong, and there is no evidence before me that it is. They are worth stating plainly because the alternative is leaving them to be discovered later, and because a superlative about an entire field currently rests on one external evaluator whose founder came out of the company being evaluated. The problem is the thinness of the evaluation layer, not the integrity of the people staffing it. A field that has decided AI models now beat physics models needs more than one live leaderboard, and it needs them run by parties with no shared history at all.
The second is a naming question that turns out to matter. Google's WeatherNext science page carries the line "How WeatherNext helped the National Hurricane Center better predict Hurricane Melissa's historic landfall in Jamaica." Melissa was October 2025. WeatherNext 3 shipped in September 2026. The model that was in the room during Melissa was an earlier member of the family. The unversioned name is accurate and it is also load-bearing: a reader who arrives on the WeatherNext 3 page reads a result that a different model earned.
Which is a pity, because the NHC's own account is more specific than the marketing and more flattering in a way Google did not claim. The report says the official NHC forecast "outperformed nearly all available model guidance at all forecast lead times," with one exception: the Google DeepMind ensemble mean, which "performed remarkably well at all time periods and was considerably more skillful than most of the other track guidance." The NHC forecasts were comparable to it through 96 hours and slightly better at 120 hours. On intensity, DeepMind's model was among the guidance that "increased forecaster confidence in predicting RI at unusually long lead times," and the human forecasters still beat every model through 96 hours. The machine was the best single instrument on the desk. The desk was better than the instrument.
Google reports the precipitation improvement as a Continuous Ranked Probability Score gain of "up to 60% against IMERG, 30% for MRMS, and 10% against rain gauge measurements for early lead times."
CRPS is a scoring rule for probabilistic forecasts. It measures the distance between the whole probability distribution a model produces and the single value that actually occurred, so lower is better and an "improvement" is a reduction in that distance. Two things about the quoted sentence: "for early lead times" governs all three figures, so none of them is a claim about the medium range; and it is not clear from Google's phrasing whether "up to" governs all three or only the first. Read strictly, these are ceilings rather than averages. That distinction is worth holding, because the argument that follows lives in the ordering of the three numbers rather than in their absolute size, and the ordering survives either reading.
Read the ladder from the top down and notice what each rung is measuring against.
IMERG is NASA's Integrated Multi-satellitE Retrievals for GPM. It combines readings from a constellation of satellites to estimate rainfall across most of the planet's surface, and NASA describes it as "particularly valuable over areas of Earth's surface that lack ground-based precipitation-measuring instruments, including oceans and remote areas." It is an inference about rain, assembled from orbit, and NASA notes in the same documentation that "precipitation in general is less certain in mountainous terrain." MRMS is NOAA's Multi-Radar Multi-Sensor system, which fuses the American radar network with surface observations, model fields and climatology into high-resolution mosaics. It is closer to the ground and much more directly measured, over a much smaller domain.
A rain gauge is a container in a field with water in it.
Satellite inference: 60 percent better. Radar mosaic: 30 percent. Bucket: 10 percent.
The improvement falls by a factor of six as the reference moves from an estimate made from orbit toward the thing a person standing outside in the rain actually experiences. The last mile here is literal. It is a distance you could walk, from a satellite's guess down to the water in the container.
Google does not explain the ordering, and there are at least two honest readings of it. The first is proximity to training signal: the model learned precipitation from IMERG and from a satellite-radar reanalysis, so it should be expected to match IMERG best and an independent instrument worst. On that reading the ladder is exactly what it looks like, a frontier that recedes as you approach the ground.
The second reading is more corrosive, and it is the one worth following. A gauge samples a point. A model predicts an area. Those are not the same quantity, and no forecast can agree perfectly with an instrument that is measuring something slightly different. Part of that 10 percent ceiling may be metrology rather than meteorology. Which would mean the ladder is partly an artifact of instrument mismatch, not evidence of anything about the model at all.
Follow that one step further and it stops being an objection. If the last mile is partly a measurement problem, then the question to ask of any forecasting system is not how good is the model but what is the gauge, and has anyone built one. Weather can ask that question because it knows what its bucket is. It has spent a century putting containers in fields and arguing about where to put them.
None of which makes a 10 percent gain against gauges a poor result. Precipitation is the hardest common field variable in the discipline, and a 10 percent CRPS improvement at early lead times is a genuine advance that will show up in people's days. What carries information is the ordering, and what the ordering suggests about where the resistance lives.
Sixty-three years ago Edward Lorenz published "Deterministic Nonperiodic Flow" in the Journal of the Atmospheric Sciences, the paper establishing that a fully deterministic system can be practically unpredictable because tiny differences in starting conditions grow without bound. Six years after that, in Tellus, he published the paper that matters more here.
"The predictability of a flow which possesses many scales of motion" proposes that two states of a fluid system differing initially by a small observational error "will evolve into two states differing as greatly as randomly chosen states of the system within a finite time interval, which cannot be lengthened by reducing the amplitude of the initial error."
That final clause is the whole thing. Better instruments do not extend the interval. The horizon is a property of the atmosphere, not of the observing system pointed at it.
The result Lorenz reaches is more interesting than a single wall, and it is the part that usually gets lost when the idea is summarised. He found that "each scale of motion possesses an intrinsic finite range of predictability," and he tabulated them. With his chosen constants, "cumulus-scale" motions can be predicted about one hour ahead, "synoptic-scale" motions a few days, and the largest scales a few weeks.
The horizon is a staircase, and it descends toward the small.
Which means the frontier WeatherNext 3 is pushing into (hourly updates, 5-kilometre grids, local convective rain) sits on the shortest steps of that staircase. Every improvement in this release points inward: sharper, faster, more local. None of them points further out. As a commercial judgment that is sound, because the money in forecasting sits at exactly those scales. On Lorenz's account it is also the direction in which the ceiling sits lowest.
Modern work has put numbers on it. In 2019, Fuqing Zhang and colleagues published "What Is the Predictability Limit of Midlatitude Weather?" in the Journal of the Atmospheric Sciences, using the 9-kilometre ECMWF operational model and identical-twin experiments with the 3-kilometre US Next-Generation Global Prediction System. Their finding is that the limit "may indeed exist and is intrinsic to the underlying dynamical system and instabilities even if the forecast model and the initial conditions are nearly perfect." They put the current practical skill horizon at around 10 days. They calculate that cutting today's initial-condition uncertainty by a full order of magnitude would extend deterministic forecasts "by up to 5 days." In the body of the paper they describe a complete loss of predictability at 15 days, and conclude that the intrinsic limit "may not be extended beyond 2 weeks."
And then the clause that should be read twice, from the same abstract: this additional headroom comes "with much less scope for improving prediction of small-scale phenomena like thunderstorms."
Perfect observations and a perfect model buy roughly five more days at synoptic scale, and considerably less than that at the scale of the storm cell sitting over one parish of Jamaica.
So there are two last miles here, and the useful thing about weather is that you can tell them apart with published numbers.
The first is an engineering last mile. It is brutal, it is expensive, and it yields. Zhang and colleagues quantify it: tropical cyclone track forecasting at the National Hurricane Center "has gained almost a day lead time per decade," such that the average five-day track error in the Atlantic in 2016 was smaller than the average two-day error in 1990. Three decades of satellites, aircraft reconnaissance, data assimilation and compute bought three days of warning. That is what the Melissa track forecast is made of, and it is the reason an island had a hurricane warning three days before the strongest landfall in its recorded history. Nothing about it was cheap or fast, and all of it worked.
The second is better described as a horizon, and it has not moved since 1969. No quantity of capital, parameters, satellites or capability crosses it, because Lorenz's argument is that the interval "cannot be lengthened by reducing the amplitude of the initial error," and fifty-seven years of subsequent work, including a 2019 study run on hardware Lorenz could not have imagined, keeps landing on roughly the same fortnight.
This is what the familiar maxims are groping at and cannot express. "The last 10 percent is more than 50 percent of the work" describes a single curve with a steepening cost. Amara's Law, that we overestimate a technology in the short run and underestimate it in the long run, describes a single curve with a lag. Both are true often enough to be quotable, and neither can tell you which curve you are standing on. Weather has two curves running simultaneously inside the same storm. The track is one. The rainfall is the other. Melissa demonstrated both in the same week, and the NHC verified both in the same report, 87 pages that include the forecasts that were wrong.
For AI in general, none of that exists. There is no Lorenz 1969 for language models. Nobody has published a scale-by-scale table of what is intrinsically predictable and what is not, no equivalent of "cumulus-scale about one hour, synoptic-scale a few days," no independent verification archive kept for thirty years, no agency releasing a post-mortem on the forecasts it got wrong. And for the benchmarks most often cited, the relationship between the evaluation set and the training data is not independently established. That is the IMERG problem, unresolved, sitting at the centre of how the field grades itself.
So when someone says the remaining distance to general intelligence is two years of engineering, and someone else says it is a wall nobody will cross, both are making a claim of the kind that weather forecasting needed sixty years and a great deal of unglamorous instrumentation to be able to make about itself. The honest position is that the measurement has not been taken. Whether it can be taken is a further question nobody has seriously specified: weather has governing equations, a physical state space and scales of motion that can be enumerated, which is what let Lorenz tabulate predictability by scale at all. What the equivalent would even be for a language model is unwritten.
DeepMind put the caveat in its own launch post: "The atmosphere will always retain a degree of unpredictability." In meteorology that is a measured claim. Someone established the degree, published it, and let others check it at resolutions the original work could not reach. Said about the rest of AI today, the same sentence would be an intuition that nobody has yet taken the trouble to measure.
Begin the conversation