WeatherNext Performance Benchmarks: RMSE, CRPS, and Hardware Comparisons Explained

WeatherNext delivers state-of-the-art medium-range weather forecasting with unbiased RMSE of ~1.10–1.20 K for 2-meter temperature on 6-day forecasts, evaluated on the WeatherBench2 ERA5 test set using standard meteorological metrics.

The google-deepmind/weathernext repository implements two distinct model families for operational weather prediction: WeatherNext 2 (Forecast Graph Neural Network) and WeatherNext 1 Gen (the GenCast diffusion-based ensemble). Both are rigorously benchmarked against ERA5 reanalysis data using industry-standard metrics. This article breaks down the exact performance numbers, how they're computed, and the architectural decisions that enable these results.

WeatherNext 2 (FGN) Benchmarks

WeatherNext 2 represents the latest generation, built on an icosahedral mesh architecture with typed-graph transformers.

Core Performance Metrics

At 0.25° resolution on the WeatherBench2 test split (2019–2022), WeatherNext 2 achieves:

  • Unbiased ensemble-mean RMSE ≈ 1.20 K for 2-meter temperature — approximately 15% improvement over persistence baseline
  • Unbiased CRPS ≈ 0.90 K — approximately 20% improvement over baseline

These figures are reported in the repository's README at lines 83–84 and derive from the Nature paper on WeatherNext.

Architecture Impact on Performance

The performance gains stem from architectural choices in weathernext/weathernext2/architecture.py, which defines mesh-aware transformer layers that reduce quadratic attention overhead. The icosahedral mesh structure, implemented in utils/typed_graph_net.py and utils/mesh_transformer.py, enables high-resolution forecasts without the memory penalties of dense global attention.

WeatherNext 1 Gen (GenCast) Benchmarks

The GenCast family uses denoising diffusion for calibrated ensemble generation, with multiple checkpoint variants.

Full-Resolution GenCast (0.25°)

The GenCast 0.25° (2019) checkpoint achieves superior deterministic accuracy:

  • Ensemble-mean RMSE ≈ 1.10 K for 2-meter temperature
  • Unbiased CRPS ≈ 0.85 K

These results appear in the GenCast-specific README at lines 35–37 and are detailed in arXiv:2312.15796.

GenCast Mini (1°) Limitations

The GenCast Mini lightweight variant at 1° resolution provides "reasonable" scores but is explicitly not representative of full-model performance. The repository documentation at lines 35–38 warns against direct comparison, as this 8-member ensemble uses simplified architecture for faster experimentation.

Hardware Impact on WeatherNext Performance

Inference hardware and attention implementation measurably affect benchmark results.

TPU vs. GPU Performance Degradation

According to docs/weathernext1_gen/cloud_vm_setup.md lines 79–84:

Configuration RMSE Impact CRPS Impact
TPU v4 with splash attention Baseline Baseline
H100 GPU with triblockdiag-MHA +~0.3% +~0.4%

The splash attention kernel is optimized for TPUs via JAX XLA. When running on GPU, users must switch to triblockdiag_mha attention, incurring minor metric degradation. This trade-off is documented for cloud VM deployments where TPU availability varies.

How WeatherNext Benchmarks Are Computed

Understanding the evaluation pipeline ensures reproducible results.

Dataset and Split

All models follow the same data protocol defined in the main README lines 55–62:

  • Training: ERA5 reanalysis, 1979–2018
  • Test evaluation: Official WeatherBench2 test split, 2019–2022

Metric Computation

RMSE is calculated per-variable after climatology detrending. The unbiased_rmse function in weathernext/utils/model_utils.py handles this computation, ensuring fair comparison across initialization times.

CRPS uses an unbiased formulation to account for varying ensemble sizes — critical since GenCast Mini (8 members) runs against full GenCast (50 members). The unbiased_crps implementation in the same module corrects for ensemble size effects.

Unified Rollout Pipeline

The weathernext/utils/rollout.py module provides autoregressive_rollout(), the single evaluation path used for all benchmarks. This guarantees metric comparability across model families. The rollout handles:

  • 6-hour autoregressive stepping
  • Consistent sampler initialization
  • State management for multi-day forecasts

Running WeatherNext Benchmark Evaluations

WeatherNext 2 Evaluation Example

import weathernext.weathernext2 as wn2
import weathernext.utils.data_utils as dutils
import weathernext.utils.rollout as rollout
import weathernext.utils.model_utils as mutils

# 1️⃣ Load pre-trained weights (see README for bucket URLs)

weights_path = "gs://dm_graphcast/WeatherNext2_2025.npz"
model = wn2.FGN.load_from_weights(weights_path)

# 2️⃣ Load WeatherBench2 test data

test_ds = dutils.load_weatherbench2(split="test", years=[2019, 2020, 2021, 2022])

# 3️⃣ Run 6-day autoregressive rollout (144 hours at 6-hour steps)

preds = rollout.autoregressive_rollout(
    model=model,
    init_state=test_ds["initial_condition"],
    horizon_steps=144 // 6,
    sampler=rollout.DefaultSampler(),
)

# 4️⃣ Compute unbiased RMSE for 2-meter temperature

rmse = mutils.unbiased_rmse(
    preds["t2m"], test_ds["target"]["t2m"]
)
print(f"Unbiased RMSE (2-m temperature): {rmse:.3f} K")

GenCast Evaluation Example

from weathernext.weathernext1_gen import gencast, denoiser
from weathernext.utils import data_utils, rollout, losses

# Load GenCast checkpoint

gc = gencast.GenCast.load_from_weights(
    "gs://dm_graphcast/GenCast_0p25deg_2019.npz"
)

# Initialize denoiser for diffusion-based generation

denoise = denoiser.Denoiser(model=gc)

# Rollout with DPM-Solver++(2S) sampler

forecast = rollout.autoregressive_rollout(
    model=denoise,
    init_state=data_utils.load_initial_condition(),
    horizon_steps=24,  # 6-day forecast at 6-hour steps

    sampler=rollout.DPMSolverPP2S(),
)

# Compute unbiased CRPS across all variables

crps = losses.unbiased_crps(forecast, data_utils.load_targets())
print(f"Unbiased CRPS: {crps:.3f} K")

Key Source Files for Benchmark Verification

File Purpose
weathernext/weathernext2/architecture.py FGN mesh-aware transformer layers
weathernext/weathernext1_gen/gencast.py GenCast model builder and sampler connector
weathernext/utils/rollout.py Unified autoregressive evaluation pipeline
weathernext/utils/model_utils.py unbiased_rmse() and unbiased_crps() implementations
docs/weathernext1_gen/cloud_vm_setup.md Hardware-specific performance notes
README.md High-level benchmark expectations and weight locations

Summary

  • WeatherNext 2 (FGN) achieves ~1.20 K unbiased RMSE and ~0.90 K CRPS at 0.25° resolution, with 15–20% improvements over persistence baselines
  • GenCast 0.25° reaches ~1.10 K RMSE and ~0.85 K CRPS through diffusion-based ensemble generation
  • GenCast Mini (1°) provides faster experimentation but is not benchmark-comparable to full-resolution models
  • Hardware matters: TPU v4 with splash attention yields optimal scores; GPU with triblockdiag_mha incurs ~0.3–0.4% metric degradation
  • Unified evaluation: The rollout.py pipeline and model_utils.py metrics ensure fair comparison across all model variants on WeatherBench2 ERA5 data

Frequently Asked Questions

What dataset is used for WeatherNext performance benchmarks?

All benchmarks use the WeatherBench2 test split of ERA5 reanalysis data, covering 2019–2022. Models are trained on 1979–2018 data. This standardization allows direct comparison with other published weather models.

Why are there two model families with different performance numbers?

WeatherNext 2 (FGN) and WeatherNext 1 Gen (GenCast) represent fundamentally different approaches: FGN uses graph neural networks with deterministic forecasting, while GenCast employs denoising diffusion for explicit ensemble generation. GenCast achieves slightly better RMSE (~1.10 K vs. ~1.20 K) but requires more complex inference; FGN offers simpler deployment with competitive probabilistic scores.

How does hardware choice affect my WeatherNext benchmark results?

Running on TPU v4 with the default splash attention kernel produces the reference scores reported in papers. Switching to H100 GPU requires triblockdiag_mha attention, which degrades RMSE by ~0.3% and CRPS by ~0.4% due to implementation differences. For publication-quality benchmarks, TPU execution is strongly recommended.

Can I reproduce the exact CRPS numbers if I use fewer ensemble members?

The unbiased_crps() function in model_utils.py applies corrections for ensemble size, so scores remain comparable whether using GenCast Mini (8 members) or full GenCast (50 members). However, the Mini model's architecture differences mean its absolute performance should not be directly compared to 0.25° checkpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →