Generated by Codex with GPT 5.6 Sol XHigh

The hardest problem in operational weather prediction is not merely generating a plausible future. It is keeping that prediction anchored to a fast-changing atmosphere while preserving both global coherence and the local detail people actually experience. WeatherNext 3 attacks those requirements together by changing the model’s data contract: it learns from live observations, refreshes hourly, and emits several kinds of forecast from one shared system.

The official Google DeepMind blog published the post on September 3, 2026, with work from Google DeepMind and Google Research. Its central contribution is an end-to-end global forecasting model that combines hourly geostationary-satellite mosaics with conventional historical weather analysis in a Functional Generative Network mesh transformer. Instead of producing only one uniform grid, the system natively predicts dense atmospheric fields, discrete cyclone tracks, and values at sparse weather-station coordinates.

One model across several scales

Previous WeatherNext forecasts used a 25-kilometer grid at six-hour intervals. WeatherNext 3 produces a new forecast every hour and varies spatial resolution according to the variable: temperature and moisture near the surface can reach 5 kilometers, other surface quantities 10 kilometers, and atmospheric variables such as wind speed 25 kilometers. That is roughly a fivefold increase in the finest spatial detail while also increasing forecast frequency by a factor of six.

The choice to use multiple native resolutions matters. A single coarse global grid smooths away coastlines, valleys, mountains, and sharp storm boundaries. Uniformly raising the resolution of every variable, however, would spend enormous compute on fields that may not benefit equally. WeatherNext 3 instead shares a common mesh-transformer representation while letting each output use the granularity appropriate to its physical behavior and operational value.

This also avoids treating every useful forecast as a post-processing problem. Station-level predictions do not have to be reconstructed later from nearby grid cells, and cyclone tracks do not have to be inferred only after generating a dense field. The model can express continuous fields, point observations, and structured trajectories through the same forecasting core. That makes the architecture closer to a multi-product prediction platform than a model built around one canonical tensor.

The post says the higher-resolution temperature output resolves local UK topography that WeatherNext 2 blurred into blocky regions. The important engineering point is broader than the visual improvement: downstream consumers no longer need to recover detail that the upstream representation discarded. Preserving the needed resolution at the model boundary reduces both information loss and the amount of compensating logic required later.

Moving observations into the learning loop

Most AI weather models are initialized and trained with analysis data produced by numerical weather prediction systems. Those analyses are useful, carefully constructed estimates of atmospheric state, but the physics-based assimilation pipeline introduces roughly six hours of delay and can carry its own biases into the learned model. Fast-changing rain, clouds, and surface temperature can move materially during that gap.

WeatherNext 3 adds live mosaics from geostationary satellites, giving the model a continuously updated view of cloud systems and other atmospheric signals. It also trains directly on sparse weather-station observations so that local measurements and topography can influence 5-kilometer surface forecasts. The conventional analysis is still present; the system augments it rather than pretending that raw sensors alone provide a complete atmospheric state.

This is a consequential systems decision. The model must reconcile inputs with very different geometries, coverage, noise, and update rates: global gridded analyses, image-like satellite mosaics, and irregular point measurements. The mesh architecture becomes an integration layer as much as a predictor. By absorbing these sources directly, the system shortens the path between a new observation and a new forecast, enabling an hourly operational cadence instead of waiting for another full analysis cycle.

The same design changes where errors can enter. A model trained only to reproduce analysis data may become excellent at imitating the assumptions of the upstream simulator. Training against real observations gives the learner a route to correct those inherited biases, though it also makes data quality, calibration, missing coverage, and sensor drift part of the production model’s responsibility. Fresher inputs are valuable only if the ingest and validation pipeline is reliable enough to support them continuously.

Treating precipitation as its own measurement problem

Rain and snow are particularly difficult because they emerge from small, rapidly evolving cloud processes and are easy for coarse models to smear across space. WeatherNext 3 trains on NASA’s IMERG satellite precipitation product and Google’s global precipitation reanalysis derived from satellite radar rather than relying on one generic target for every weather variable.

Google reports improvements in Continuous Ranked Probability Score of up to 60% against IMERG, 30% against the U.S. Multi-Radar/Multi-Sensor system, and 10% against rain gauges at early lead times. CRPS evaluates a probabilistic forecast by rewarding both calibration and sharpness, so the result is more informative than measuring only whether the model’s average prediction was close. The range across reference datasets is also a useful caution: forecast quality depends on which observation system is treated as truth, and no single number captures every geography or lead time.

The production-facing result is sharper storm structure and better medium-range probabilities, not simply prettier high-resolution maps. Google says forecasts for precipitation a day or more ahead become up to 50% more accurate in its consumer products, with the largest gains in historically underserved regions. The system also adds variables designed for renewable-energy planning, including wind at roughly turbine height, cloud cover, and surface solar radiation. Those outputs illustrate why architecture should begin with decisions users need to make, then preserve the temporal and spatial fidelity those decisions require.

The broader engineering lesson

WeatherNext 3 is being deployed as both a model and a data service. Google is integrating it into Search, Gemini, Maps, the Google Maps Platform Weather API, and Earth Engine, while exposing forecast data through BigQuery and Cloud Storage. Serving precomputed, hourly global fields lets many consumers share the expensive forecasting run and query the representation they need, instead of each team operating a specialized model.

The post does not disclose every training, compute, latency, or reliability detail, and its strongest accuracy figures are improvements of up to a stated maximum rather than guarantees across all conditions. Google also directs users to national meteorological agencies for official warnings. Those limits matter because an operational forecast is a safety-relevant pipeline: model skill, sensor availability, update deadlines, uncertainty communication, and graceful degradation all determine whether a technically better forecast is trustworthy in practice.

The transferable lesson is that a model-system breakthrough can come from removing boundaries between stages, not only from scaling the predictor. WeatherNext 3 connects fresher observations, a shared multi-resolution representation, variable-specific supervision, structured outputs, and a reusable serving layer. That co-design reduces stale inputs and lossy handoffs while keeping one coherent forecast behind many products. For any real-time ML platform, the same question is worth asking: which limitations belong to the model, and which are artifacts of an upstream data pipeline or downstream reconstruction step that the model could safely absorb?