WeatherNext Evaluation Metrics: CRPS, MAE, and MAD Explained

WeatherNext uses the Continuous Ranked Probability Score (CRPS) as its primary evaluation metric, computed from Mean Absolute Error (MAE) and Mean Absolute Difference (MAD) to assess probabilistic forecast quality.

WeatherNext, developed by Google DeepMind, includes both the WeatherNext 2 global forecast model and the WeatherNext Cyclones specialized model. Understanding the evaluation metrics for WeatherNext is essential for interpreting model performance and reproducing results. This article breaks down how CRPS is implemented in the codebase, where to find the relevant functions, and how to apply them to your own ensemble predictions.

What is CRPS in WeatherNext?

The Continuous Ranked Probability Score (CRPS) measures how well a probabilistic forecast matches observed outcomes. Unlike deterministic metrics, CRPS evaluates the full ensemble distribution, rewarding both accuracy and appropriate uncertainty quantification.

In weathernext/weathernext2/fgn.py, the FGN class implements crps_loss, which returns the scalar CRPS value plus diagnostic breakdowns. The implementation follows this formula:


CRPS = MAE – 0.5 * MAD

CRPS Components: MAE and MAD

WeatherNext's CRPS decomposes into two interpretable quantities that are also exposed as standalone diagnostics.

Mean Absolute Error (MAE)

MAE measures the average absolute difference between each ensemble member's prediction and the target. In WeatherNext, targets are shift-corrected to account for systematic biases before computing this error.

Mean Absolute Difference (MAD)

MAD captures ensemble spread by computing the average absolute difference between all pairs of ensemble members. Higher MAD indicates more diverse predictions; near-zero MAD suggests the ensemble has collapsed to a single solution.

Biased vs. Unbiased Estimators

The crps_loss method supports two variance estimators via the unbiased flag:

  • unbiased=True (default): Normalizes MAD by n · (n‑1), providing an unbiased estimator for finite ensembles
  • unbiased=False: Normalizes by n², the biased but simpler estimator

This choice affects how ensemble size influences the metric, particularly for small ensembles.

How to Compute CRPS in WeatherNext

Use the fgn module to calculate CRPS on your own ensemble predictions:

import xarray as xr
from weathernext.weathernext2 import fgn

# preds: xarray DataArray with ensemble dimension "sample"

# targets: xarray DataArray with identical coordinates (excluding "sample")

crps, crps_diag = fgn.FGN().crps_loss(
    all_predictions=preds,
    targets=targets,
    unbiased=True,  # use unbiased estimator

)

print("CRPS loss:", crps)
print("Per-variable diagnostics:", crps_diag)

The crps_diag output contains per-variable MAE and MAD values, enabling detailed analysis of which atmospheric variables contribute most to forecast error.

Visualizing WeatherNext Evaluation Metrics

The docs/weathernext2/wn2_demo.ipynb notebook demonstrates practical CRPS usage. Key visualizations include:

  • Ensemble-mean fields — deterministic view of the probabilistic forecast
  • CRPS spatial maps — geographic distribution of forecast skill
  • Per-variable diagnostics — breakdown by pressure level and variable type

These visualizations help identify regions and variables where WeatherNext excels or struggles.

Supporting Infrastructure for Evaluation

Several utility modules support WeatherNext's evaluation pipeline:

File Purpose
weathernext/utils/losses.py Generic loss functions consumed by crps_loss
weathernext/utils/rollout.py Autoregressive rollout for multi-step evaluation
weathernext/weathernext2/fgn.py Core CRPS, MAE, and MAD implementation

The rollout utilities enable evaluation beyond single time steps, chaining predictions through the model's autoregressive mechanism to assess long-range forecast degradation.

Summary

  • Primary WeatherNext evaluation metric: Continuous Ranked Probability Score (CRPS), implemented in fgn.FGN().crps_loss
  • Core components: Mean Absolute Error (MAE) for accuracy and Mean Absolute Difference (MAD) for ensemble spread
  • Estimator variants: Biased (÷ n²) and unbiased (÷ n·(n‑1)) via the unbiased parameter
  • Diagnostic outputs: Per-variable MAE and MAD for targeted analysis
  • Practical entry points: weathernext2/fgn.py for computation, wn2_demo.ipynb for visualization

Frequently Asked Questions

What does CRPS measure that RMSE does not?

CRPS evaluates the full probability distribution of an ensemble forecast, while RMSE only measures deterministic point forecasts. This makes CRPS sensitive to both forecast accuracy and appropriate uncertainty quantification—an overconfident ensemble with collapsed spread receives worse CRPS than a properly dispersed one, even if the ensemble mean is identical.

How does the unbiased CRPS estimator differ from the biased version?

The unbiased estimator divides MAD by n·(n‑1) rather than n², correcting for finite-sample bias in ensemble variance estimation. This matters most for small ensembles (n < 10), where the biased estimator systematically underestimates true ensemble spread. WeatherNext uses unbiased=True by default.

Can I extract MAE and MAD separately from the CRPS calculation?

Yes. The crps_loss method returns a tuple where the second element (crps_diag) contains per-variable MAE and MAD diagnostics. These values are computed as intermediate steps in the CRPS formula, so retrieving them adds no computational overhead.

Where is the WeatherNext CRPS implementation located?

The core implementation resides in weathernext/weathernext2/fgn.py within the FGN.crps_loss method. This file belongs to the WeatherNext 2 model; cyclone-specific evaluation may use additional metrics not covered in the base implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →