# WeatherNext Input Requirements: Complete Guide to Data Format, Resolution, and Variables

> Understand WeatherNext input requirements. Learn about Zarr datasets, ERA5/HRES fields, resolution, and variables needed for optimal model performance.

- Repository: [Google DeepMind/weathernext](https://github.com/google-deepmind/weathernext)
- Tags: tutorial
- Published: 2026-08-16

---

**WeatherNext requires Zarr-based xarray datasets containing ERA5 or HRES atmospheric fields at 0.25° resolution with specific pressure-level variables, which are then normalized using `weathernext.utils.data_utils` before model ingestion.**

WeatherNext is a research-grade global medium-range forecasting system developed by Google DeepMind. Understanding its input requirements is essential for anyone running inference or training custom variants. This guide breaks down the exact data specifications, file formats, and preprocessing steps needed to feed atmospheric data into the model.

## Data Sources for WeatherNext

WeatherNext accepts input from two primary sources that follow ECMWF conventions.

### ERA5 Reanalysis Data

The **ERA5** dataset from the Copernicus Climate Change Service serves as the standard training data. It provides consistent, long-term atmospheric reanalysis spanning decades of historical weather. Access ERA5 through the **WeatherBench 2** data portal, which serves data as optimized Zarr stores.

### HRES Operational Forecasts

For operational inference and fine-tuning, WeatherNext uses **HRES** (High-Resolution Ensemble) data from ECMWF's operational forecasting system. HRES provides the initial conditions for real-time medium-range forecasts at matching resolution and variable specifications.

Both sources are typically accessed as **Zarr** stores and loaded into **xarray** `Dataset` objects for processing. According to the repository README (lines 55-63, 71-73), the data pipeline expects these formats to ensure efficient, chunked access to global atmospheric fields.

## Spatial and Temporal Specifications

WeatherNext has strict requirements for grid geometry and initialization timing.

### Spatial Resolution

| Checkpoint | Resolution | Approximate Grid Spacing |
|------------|-----------|-------------------------|
| WeatherNext 2 (WN2) default | 0.25° | ~30 km |
| Mini checkpoints | 1.0° | ~111 km |

The grid layout must match the model's internal mesh structure—either icosahedral or regular latitude-longitude depending on the checkpoint. Mismatched resolutions require regridding before ingestion.

### Vertical Levels

Input fields must include the standard set of **pressure levels** used during training:

- Full range: 1000 hPa down to 100 hPa
- Typical variables per level: geopotential height, temperature, u-wind component, v-wind component

Surface-only variables (e.g., 2-meter temperature, 10-meter winds) are provided separately without vertical coordinates.

### Temporal Structure

Forecasts initialize from a **single analysis time**—the initial condition. The input dataset contains one time step (or a minimal window for spin-up protocols) that seeds the model's **autoregressive rollout**. Each subsequent forecast step becomes the input for the next prediction, propagating forward in 6-hour increments.

## Required Variables and Normalization

The `weathernext.utils.data_utils` module defines the complete variable set and provides normalization utilities.

### Core Atmospheric Variables

- **`t2m`** — 2-meter temperature
- **`u10`, `v10`** — 10-meter wind components (zonal and meridional)
- **`geopotential`** — Geopotential height on pressure levels
- **`surface_pressure`** — Surface pressure field

Additional pressure-level fields include temperature (`t`), u-wind (`u`), and v-wind (`v`) across the vertical stack.

### Normalization Pipeline

Raw meteorological values must be scaled to the ranges expected by the model's neural network. The `data_utils.normalize_inputs()` function applies pre-computed statistics derived from the training distribution. This step is **mandatory**—un-normalized inputs produce invalid forecasts.

## Loading and Preprocessing Code Examples

### Inference with HRES Initial Conditions

```python
import xarray as xr
from weathernext.utils import data_utils

# Path to HRES Zarr store (WeatherBench2 format)

zarr_path = "gs://weatherbench2/hres/2024-01-01.zarr"

# Load dataset with required variables

ds = xr.open_zarr(zarr_path, consolidated=True)

# Normalize to model's internal scale

inputs = data_utils.normalize_inputs(ds)

# Initialize WeatherNext2 model

from weathernext.weathernext2.fgn import WeatherNext2
model = WeatherNext2(checkpoint="WeatherNext2_<2025_model1>.npz")

# Execute 10-day autoregressive forecast (240 steps × 6 hours)

forecast = model.auto_regressive_rollout(inputs, rollout_steps=240)

```

### Training Data Preparation with ERA5

```python
import xarray as xr
from weathernext.utils import data_utils

# ERA5 Zarr store at 0.25° resolution

zarr_path = "gs://weatherbench2/era5/2023-06.zarr"

# Load and select specific variables

ds = xr.open_zarr(zarr_path, consolidated=True)
train_inputs = data_utils.select_and_normalise(
    ds,
    variables=["t2m", "u10", "v10", "geopotential", "surface_pressure"]
)

# Feed to training loop in weathernext.weathernext2.fgn.train

```

## Key Implementation Files

Understanding the codebase structure helps debug input-related issues:

- **[`weathernext/utils/data_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/data_utils.py)** — Dataset loading, variable selection, and normalization functions
- **[`weathernext/utils/model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/model_utils.py)** — Model weight loading and mesh configuration
- **[`weathernext/weathernext2/fgn.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext2/fgn.py)** — Core `WeatherNext2` class with `auto_regressive_rollout()` method
- **`docs/weathernext2/wn2_demo.ipynb`** — Interactive demonstration of complete input-to-forecast pipeline

## Summary

WeatherNext input requirements center on four pillars:

- **Format**: Zarr-based xarray datasets from ERA5 (training) or HRES (inference)
- **Resolution**: 0.25° for standard checkpoints, 1.0° for mini variants
- **Variables**: Complete set of surface and pressure-level fields defined in `data_utils`
- **Preprocessing**: Mandatory normalization via `weathernext.utils.data_utils` before model ingestion

Failure at any stage—incorrect resolution, missing variables, or skipped normalization—will prevent successful model initialization or produce non-physical forecasts.

## Frequently Asked Questions

### What file format does WeatherNext require for input data?

WeatherNext requires **Zarr** format, loaded as **xarray** `Dataset` objects. Zarr provides chunked, compressed storage optimized for cloud access and parallel I/O. The WeatherBench 2 portal serves both ERA5 and HRES data in this format, and `xr.open_zarr()` is the standard entry point in the preprocessing pipeline.

### Can I use my own observational data instead of ERA5 or HRES?

Custom data is possible but requires strict adherence to WeatherNext conventions. Your dataset must match the target resolution, include all required variables on compatible vertical levels, and follow the same coordinate naming (latitude, longitude, level, time). You must also compute normalization statistics matching the model's training distribution, or re-train the model on your data source.

### What happens if my input resolution doesn't match 0.25°?

Resolution mismatches require explicit **regridding** before ingestion. The model does not perform on-the-fly interpolation. Use `xarray` regridding utilities or specialized meteorological packages (e.g., `xesmf`, `metpy`) to interpolate your data to the checkpoint's expected grid. Running at incorrect resolution without regridding will raise errors or produce spatially aliased forecasts.

### Do I need internet access to run WeatherNext inference?

Once input data is staged locally, no internet connection is required. However, accessing the canonical data sources (Google Cloud Storage buckets for WeatherBench 2) requires connectivity. For air-gapped environments, download Zarr stores in full using tools like `gsutil`, then point `xr.open_zarr()` to the local path.