# WeatherNext Performance Benchmarks: RMSE, CRPS, and Hardware Comparisons Explained

> Discover WeatherNext performance benchmarks including RMSE and CRPS. See hardware comparisons and learn how WeatherNext achieves state-of-the-art weather forecasting. Get unbiased results.

- Repository: [Google DeepMind/weathernext](https://github.com/google-deepmind/weathernext)
- Tags: performance
- Published: 2026-08-16

---

**WeatherNext delivers state-of-the-art medium-range weather forecasting with unbiased RMSE of ~1.10–1.20 K for 2-meter temperature on 6-day forecasts, evaluated on the WeatherBench2 ERA5 test set using standard meteorological metrics.**

The google-deepmind/weathernext repository implements two distinct model families for operational weather prediction: **WeatherNext 2** (Forecast Graph Neural Network) and **WeatherNext 1 Gen** (the GenCast diffusion-based ensemble). Both are rigorously benchmarked against ERA5 reanalysis data using industry-standard metrics. This article breaks down the exact performance numbers, how they're computed, and the architectural decisions that enable these results.

## WeatherNext 2 (FGN) Benchmarks

WeatherNext 2 represents the latest generation, built on an icosahedral mesh architecture with typed-graph transformers.

### Core Performance Metrics

At **0.25° resolution** on the WeatherBench2 test split (2019–2022), WeatherNext 2 achieves:

- **Unbiased ensemble-mean RMSE ≈ 1.20 K** for 2-meter temperature — approximately 15% improvement over persistence baseline
- **Unbiased CRPS ≈ 0.90 K** — approximately 20% improvement over baseline

These figures are reported in the repository's README at lines [83–84](https://github.com/google-deepmind/weathernext/blob/main/README.md#L83-L84) and derive from the Nature paper on WeatherNext.

### Architecture Impact on Performance

The performance gains stem from architectural choices in [`weathernext/weathernext2/architecture.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext2/architecture.py), which defines mesh-aware transformer layers that reduce quadratic attention overhead. The icosahedral mesh structure, implemented in [`utils/typed_graph_net.py`](https://github.com/google-deepmind/weathernext/blob/main/utils/typed_graph_net.py) and [`utils/mesh_transformer.py`](https://github.com/google-deepmind/weathernext/blob/main/utils/mesh_transformer.py), enables high-resolution forecasts without the memory penalties of dense global attention.

## WeatherNext 1 Gen (GenCast) Benchmarks

The GenCast family uses denoising diffusion for calibrated ensemble generation, with multiple checkpoint variants.

### Full-Resolution GenCast (0.25°)

The **GenCast 0.25° (2019)** checkpoint achieves superior deterministic accuracy:

- **Ensemble-mean RMSE ≈ 1.10 K** for 2-meter temperature
- **Unbiased CRPS ≈ 0.85 K**

These results appear in the GenCast-specific README at lines [35–37](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/README.md#L35-L37) and are detailed in arXiv:2312.15796.

### GenCast Mini (1°) Limitations

The **GenCast Mini** lightweight variant at 1° resolution provides "reasonable" scores but is explicitly **not representative** of full-model performance. The repository documentation at lines [35–38](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/README.md#L35-L38) warns against direct comparison, as this 8-member ensemble uses simplified architecture for faster experimentation.

## Hardware Impact on WeatherNext Performance

Inference hardware and attention implementation measurably affect benchmark results.

### TPU vs. GPU Performance Degradation

According to [`docs/weathernext1_gen/cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/cloud_vm_setup.md) lines [79–84](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/cloud_vm_setup.md#L79-L84):

| Configuration | RMSE Impact | CRPS Impact |
|-------------|-------------|-------------|
| TPU v4 with `splash` attention | Baseline | Baseline |
| H100 GPU with `triblockdiag-MHA` | +~0.3% | +~0.4% |

The `splash` attention kernel is optimized for TPUs via JAX XLA. When running on GPU, users must switch to `triblockdiag_mha` attention, incurring minor metric degradation. This trade-off is documented for cloud VM deployments where TPU availability varies.

## How WeatherNext Benchmarks Are Computed

Understanding the evaluation pipeline ensures reproducible results.

### Dataset and Split

All models follow the same data protocol defined in the main README lines [55–62](https://github.com/google-deepmind/weathernext/blob/main/README.md#L55-L62):

- **Training**: ERA5 reanalysis, 1979–2018
- **Test evaluation**: Official WeatherBench2 test split, 2019–2022

### Metric Computation

**RMSE** is calculated per-variable after climatology detrending. The `unbiased_rmse` function in [`weathernext/utils/model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/model_utils.py) handles this computation, ensuring fair comparison across initialization times.

**CRPS** uses an unbiased formulation to account for varying ensemble sizes — critical since GenCast Mini (8 members) runs against full GenCast (50 members). The `unbiased_crps` implementation in the same module corrects for ensemble size effects.

### Unified Rollout Pipeline

The [`weathernext/utils/rollout.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/rollout.py) module provides `autoregressive_rollout()`, the single evaluation path used for all benchmarks. This guarantees metric comparability across model families. The rollout handles:

- 6-hour autoregressive stepping
- Consistent sampler initialization
- State management for multi-day forecasts

## Running WeatherNext Benchmark Evaluations

### WeatherNext 2 Evaluation Example

```python
import weathernext.weathernext2 as wn2
import weathernext.utils.data_utils as dutils
import weathernext.utils.rollout as rollout
import weathernext.utils.model_utils as mutils

# 1️⃣ Load pre-trained weights (see README for bucket URLs)

weights_path = "gs://dm_graphcast/WeatherNext2_2025.npz"
model = wn2.FGN.load_from_weights(weights_path)

# 2️⃣ Load WeatherBench2 test data

test_ds = dutils.load_weatherbench2(split="test", years=[2019, 2020, 2021, 2022])

# 3️⃣ Run 6-day autoregressive rollout (144 hours at 6-hour steps)

preds = rollout.autoregressive_rollout(
    model=model,
    init_state=test_ds["initial_condition"],
    horizon_steps=144 // 6,
    sampler=rollout.DefaultSampler(),
)

# 4️⃣ Compute unbiased RMSE for 2-meter temperature

rmse = mutils.unbiased_rmse(
    preds["t2m"], test_ds["target"]["t2m"]
)
print(f"Unbiased RMSE (2-m temperature): {rmse:.3f} K")

```

### GenCast Evaluation Example

```python
from weathernext.weathernext1_gen import gencast, denoiser
from weathernext.utils import data_utils, rollout, losses

# Load GenCast checkpoint

gc = gencast.GenCast.load_from_weights(
    "gs://dm_graphcast/GenCast_0p25deg_2019.npz"
)

# Initialize denoiser for diffusion-based generation

denoise = denoiser.Denoiser(model=gc)

# Rollout with DPM-Solver++(2S) sampler

forecast = rollout.autoregressive_rollout(
    model=denoise,
    init_state=data_utils.load_initial_condition(),
    horizon_steps=24,  # 6-day forecast at 6-hour steps

    sampler=rollout.DPMSolverPP2S(),
)

# Compute unbiased CRPS across all variables

crps = losses.unbiased_crps(forecast, data_utils.load_targets())
print(f"Unbiased CRPS: {crps:.3f} K")

```

## Key Source Files for Benchmark Verification

| File | Purpose |
|------|---------|
| [`weathernext/weathernext2/architecture.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext2/architecture.py) | FGN mesh-aware transformer layers |
| [`weathernext/weathernext1_gen/gencast.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext1_gen/gencast.py) | GenCast model builder and sampler connector |
| [`weathernext/utils/rollout.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/rollout.py) | Unified autoregressive evaluation pipeline |
| [`weathernext/utils/model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/model_utils.py) | `unbiased_rmse()` and `unbiased_crps()` implementations |
| [`docs/weathernext1_gen/cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/cloud_vm_setup.md) | Hardware-specific performance notes |
| [`README.md`](https://github.com/google-deepmind/weathernext/blob/main/README.md) | High-level benchmark expectations and weight locations |

## Summary

- **WeatherNext 2 (FGN)** achieves ~1.20 K unbiased RMSE and ~0.90 K CRPS at 0.25° resolution, with 15–20% improvements over persistence baselines
- **GenCast 0.25°** reaches ~1.10 K RMSE and ~0.85 K CRPS through diffusion-based ensemble generation
- **GenCast Mini (1°)** provides faster experimentation but is not benchmark-comparable to full-resolution models
- **Hardware matters**: TPU v4 with `splash` attention yields optimal scores; GPU with `triblockdiag_mha` incurs ~0.3–0.4% metric degradation
- **Unified evaluation**: The [`rollout.py`](https://github.com/google-deepmind/weathernext/blob/main/rollout.py) pipeline and [`model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/model_utils.py) metrics ensure fair comparison across all model variants on WeatherBench2 ERA5 data

## Frequently Asked Questions

### What dataset is used for WeatherNext performance benchmarks?

All benchmarks use the **WeatherBench2** test split of ERA5 reanalysis data, covering 2019–2022. Models are trained on 1979–2018 data. This standardization allows direct comparison with other published weather models.

### Why are there two model families with different performance numbers?

**WeatherNext 2 (FGN)** and **WeatherNext 1 Gen (GenCast)** represent fundamentally different approaches: FGN uses graph neural networks with deterministic forecasting, while GenCast employs denoising diffusion for explicit ensemble generation. GenCast achieves slightly better RMSE (~1.10 K vs. ~1.20 K) but requires more complex inference; FGN offers simpler deployment with competitive probabilistic scores.

### How does hardware choice affect my WeatherNext benchmark results?

Running on TPU v4 with the default `splash` attention kernel produces the reference scores reported in papers. Switching to H100 GPU requires `triblockdiag_mha` attention, which degrades RMSE by ~0.3% and CRPS by ~0.4% due to implementation differences. For publication-quality benchmarks, TPU execution is strongly recommended.

### Can I reproduce the exact CRPS numbers if I use fewer ensemble members?

The `unbiased_crps()` function in [`model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/model_utils.py) applies corrections for ensemble size, so scores remain comparable whether using GenCast Mini (8 members) or full GenCast (50 members). However, the Mini model's architecture differences mean its absolute performance should not be directly compared to 0.25° checkpoints.