# What Data Does WeatherNext Use for Training? ERA5, HRES, and IBTrACS Explained

> Discover the diverse datasets powering WeatherNext. Learn how ERA5, HRES, and IBTrACS data train advanced weather prediction models for superior accuracy.

- Repository: [Google DeepMind/weathernext](https://github.com/google-deepmind/weathernext)
- Tags: deep-dive
- Published: 2026-08-10

---

**WeatherNext models are trained on large-scale atmospheric datasets including ECMWF ERA5 reanalysis data accessed via WeatherBench2, high-resolution HRES operational forecasts from the IFS cycle, and IBTrACS historical cyclone tracks for tropical storm prediction.**

The `google-deepmind/weathernext` repository provides state-of-the-art weather forecasting models that rely on specific, publicly available atmospheric datasets for training. Understanding what data WeatherNext uses for training is essential for researchers looking to reproduce results or adapt these models for new forecasting tasks. The codebase ingests multi-decadal reanalysis and operational forecast data stored in cloud-optimized Zarr formats.

## Primary WeatherNext Training Datasets

### ERA5 Reanalysis via WeatherBench2

The foundation of WeatherNext training data comes from the **ECMWF ERA5 reanalysis dataset**. According to the repository's documentation, this data is accessed through the WeatherBench2 data-loader as a Zarr store, providing the model with a complete, globally consistent history of atmospheric variables spanning multiple decades.

### HRES Operational Forecasts (IFS Cycle 47r1)

For fine-tuning and the operational version of WeatherNext 2, the model incorporates **ECMWF HRES (high-resolution) forecast data** derived from the IFS (Integrated Forecast System) cycle 47r1. This dataset provides higher-resolution fields that the model must learn to ingest directly, bridging the gap between historical reanalysis and real-time forecasting.

### IBTrACS Cyclone Tracks

The cyclone-specific models, including WeatherNext Cyclones, utilize the **International Best Track Archive for Climate Stewardship (IBTrACS)** dataset. This provides historical tropical cyclone positions and intensities for supervised learning of track prediction, enabling specialized forecasting for extreme weather events.

## Data Loading Implementation in the WeatherNext Repository

The repository provides specific utilities for ingesting these datasets. In [`weathernext/utils/data_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/data_utils.py), functions handle the loading of ERA5 and HRES datasets along with time slicing and derived variable computation.

```python

# Load ERA5 from WeatherBench2 (Zarr)

import xarray as xr
from weathernext.utils import data_utils

era5_path = "gs://weatherbench2/era5.zarr"   # Google Cloud bucket (public)

era5_ds = xr.open_zarr(era5_path, consolidated=True)

# Add derived variables (e.g., total incident solar radiation)

data_utils.add_derived_vars(era5_ds)

# Load HRES operational forecasts for fine‑tuning

hres_path = "gs://weatherbench2/hres.zarr"
hres_ds = xr.open_zarr(hres_path, consolidated=True)

# Load IBTrACS cyclone tracks

from weathernext.cyclones import ibtracs_processing_utils as ibtracs
ibtracs_ds = ibtracs.load_ibtracs("gs://weatherbench2/ibtracs.zarr")

```

These snippets illustrate the typical data-ingestion pipeline implemented in the codebase: opening a Zarr store, enriching it with derived features via `data_utils.add_derived_vars()`, and optionally merging cyclone track information through `ibtracs.load_ibtracs()`.

## Supporting Data Utilities and Derived Variables

Beyond raw atmospheric fields, the WeatherNext training pipeline computes additional environmental features. The [`weathernext/utils/solar_radiation.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/solar_radiation.py) module calculates total incident solar radiation from ERA5 data for use as an input channel. While the repository references Copernicus climate data and NOAA products in supporting utilities, these serve mainly for derived features and evaluation rather than core model training.

## Summary

- **ERA5 Reanalysis**: The primary training data source accessed via WeatherBench2 Zarr stores, providing global atmospheric history for base model training.
- **HRES Operational Forecasts**: High-resolution IFS cycle 47r1 data used for fine-tuning WeatherNext 2 operational models to handle real-time forecast characteristics.
- **IBTrACS Dataset**: Historical tropical cyclone tracks used specifically for training cyclone prediction models in the `weathernext/cyclones` module.
- **Zarr Format**: All core datasets are stored in cloud-optimized Zarr format on Google Cloud Storage for efficient access.
- **Derived Variables**: The codebase enriches raw data with computed features like solar radiation through utilities in [`data_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/data_utils.py).

## Frequently Asked Questions

### What is the main data source for training WeatherNext models?

The primary data source is the ECMWF ERA5 reanalysis dataset accessed through WeatherBench2 as Zarr stores. This provides the global atmospheric variables necessary for training both deterministic and probabilistic forecasting models in the repository.

### How does WeatherNext handle high-resolution operational forecasts?

WeatherNext 2 uses ECMWF HRES operational forecast data from the IFS cycle 47r1 for fine-tuning. This allows the model to adapt from historical reanalysis to the characteristics of real-time high-resolution forecast fields.

### What dataset is used for cyclone-specific WeatherNext models?

The cyclone-specific variants utilize the IBTrACS (International Best Track Archive for Climate Stewardship) dataset. This supplies historical tropical cyclone positions and intensities for supervised track prediction training, loaded via [`ibtracs_processing_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/ibtracs_processing_utils.py).

### Where are the WeatherNext training datasets stored?

All major datasets are stored in Google Cloud Storage buckets (e.g., `gs://weatherbench2/`) using the Zarr format. The repository accesses these via `xarray.open_zarr()` as implemented in the data utilities, enabling efficient cloud-based data loading without local storage requirements.