Direct Cyclone Tracker Implementation and Track Extraction in WeatherNext

The DirectTracker class in weathernext/cyclones/direct_tracker.py implements a greedy cyclone-tracking algorithm that initializes cyclone centers via IBTrACS data or direct cyclogenesis, evolves centers through time using momentum-based prediction and probability-weighted refinement, and returns tracks as pandas DataFrames.

The direct cyclone tracker is a core component of the WeatherNext ensemble forecasting system developed by Google DeepMind. It transforms raw gridded probabilistic forecasts into discrete tropical cyclone trajectories without requiring external reanalysis data. This article explains how the tracker is architected and how to extract complete storm tracks from model outputs.

Direct Tracker Architecture

The DirectTracker follows a three-stage pipeline: initialization, temporal evolution, and result collection. Each stage uses geodesic distance calculations and probability-weighted operations on the sphere.

Stage 1: Cyclone Center Initialization

The tracker accepts two initialization modes controlled by the __init__ method (lines 33‑59).

  • IBTrACS-guided mode: When a historical dataframe is supplied, centers are seeded at observed locations
  • Direct cyclogenesis: The algorithm scans prob_cyclone_exists fields for high-probability blobs, ranks candidates, and prunes them using minimum separation distances

The cyclogenesis logic resides in _do_direct_cyclogenesis (lines 78‑74), which delegates candidate detection to _get_cyclogenesis_candidates (lines 70‑44) and refinement to _prune_and_refine_cyclogenesis_candidates (lines 45‑92). The refinement step uses disc_radius_cyclogenesis_refinement_km to merge nearby candidates and select probability maxima.

Stage 2: Temporal Evolution with Momentum

For each lead-time step, active cyclones are advanced through four operations in _advance_single_cyclone_active_at_lead_time_t (lines 109‑124):

  1. Momentum prediction: _get_latlon_guess_with_momentum_update (lines 76‑99) extrapolates the next position using the current velocity vector scaled by momentum_constant

  2. Tracking mode refinement: The predicted position is corrected using one of three modes operating on a disc-shaped neighborhood:

    • mean: Probability-weighted average via _get_mean_latlon_of_grid_points_within_disc
    • mode: Highest-probability grid point via _get_mode_latlon_of_grid_points_within_disc
    • mode-then-mean: Mode seeding followed by mean refinement
  3. Scalar variable extraction: Two methods extract intensity variables:

    • _get_scalar_variables_via_bilinear_interpolation (lines 96‑102)
    • _get_all_scalar_variables_within_disc for probability-weighted disc averages
  4. Dissipation check: Tracks terminate when mean existence probability falls below threshold (lines 48‑55)

Stage 3: Track Assembly

Each center update generates a row via _single_row_dataframe_from_dict (lines 83‑85). Rows accumulate in lists and concatenate through pd.concat into a complete trajectory DataFrame containing latitude, longitude, lead time, valid time, wind speed, pressure, radii, and existence probability.

Core Geospatial Operations

The tracker performs all spatial calculations on the sphere using these helpers:

Function Purpose Location
_get_bounding_box_sides_in_degrees Computes lat/lon box containing a geodesic disc of radius r km direct_tracker.py lines 87‑34
_get_is_in_radius_grid_around_latlon Boolean mask for points inside a disc around a center direct_tracker.py lines 47‑33
_average_in_three_dimensions_and_project_on_sphere Probability-weighted averaging on unit sphere direct_tracker.py lines 37‑22
slice_data_array_latlon_box_with_lon_wraparound Sub-grid extraction handling 0-360° wrap-around utils/tracker_utils.py
latlon_to_cartesian / cartesian_to_latlon Spherical coordinate conversions utils/cyclone_utils.py

Extracting Cyclone Tracks: Complete Workflow

Step 1: Load Configuration

from weathernext.cyclones.direct_tracker_6h_v1_config import get_config

cfg = get_config()  # Pre-tuned for 6-hour forecast cadence

The configuration object bundles all radii, thresholds, and the tracking mode selection.

Step 2: Instantiate Tracker

tracker = cfg.tracker_constructor(
    disc_radius_scalar_variables_km=cfg.disc_radius_scalar_variables_km,
    disc_radius_mean_latlon_km=cfg.disc_radius_mean_latlon_km,
    disc_radius_mode_latlon_km=cfg.disc_radius_mode_latlon_km,
    disc_radius_mean_probability_of_existence_km=cfg.disc_radius_mean_probability_of_existence_km,
    disc_radius_cyclogenesis_refinement_km=cfg.disc_radius_cyclogenesis_refinement_km,
    tracking_mode=cfg.tracking_mode,  # 'mean', 'mode', or 'mode_then_mean'

    momentum_constant=cfg.momentum_constant,
    min_disc_radius_between_cyclogenesis_candidates_km=cfg.min_disc_radius_between_cyclogenesis_candidates_km,
    temporal_resolution_hours=cfg.temporal_resolution_hours,
)

Step 3: Run Tracking on Gridded Forecast

track_df = tracker.track(gridded_ds)  # gridded_ds: xr.Dataset with cyclone variables

The track method is defined in the abstract base class CycloneTracker (tracker_base.py) and orchestrates the three-stage pipeline.

Step 4: Analyze Output

print(track_df.columns)

# Output: ['lat', 'lon', 'lead_time', 'valid_time', 'max_sustained_wind_speed_knots',

#          'radius_of_maximum_winds', 'prob_cyclone_exists', ...]

print(track_df[['lat', 'lon', 'max_sustained_wind_speed_knots']].head())

Key Source Files

File Responsibility
weathernext/cyclones/direct_tracker.py Main DirectTracker implementation with cyclogenesis, momentum updates, and tracking modes
weathernext/cyclones/direct_tracker_6h_v1_config.py Production configuration for 6-hourly forecasts
weathernext/cyclones/tracker_base.py Abstract CycloneTracker class defining the track API
weathernext/cyclones/tracker_utils.py Dataset slicing and interpolation utilities
weathernext/cyclones/cyclone_utils.py Geodesic distances and spherical coordinate math
weathernext/cyclones/ibtracs_processing_utils.py IBTrACS to gridded variable name mapping
weathernext/cyclones/constants.py Column name and coordinate key definitions

Summary

  • The direct cyclone tracker implements greedy trajectory extraction through probability-weighted operations on forecast grids
  • Three tracking modes (mean, mode, mode-then-mean) control how center positions are refined from gridded probabilities
  • Momentum-based prediction reduces search windows and maintains track continuity across lead times
  • Tracks extract as pandas DataFrames with full temporal evolution of position, intensity, and existence probability
  • The 6-hour configuration provides production-ready parameters without manual tuning

Frequently Asked Questions

What is the difference between mean and mode tracking modes?

Mean tracking computes the probability-weighted average of all grid points within the search disc using _average_in_three_dimensions_and_project_on_sphere, producing smooth trajectories. Mode tracking selects the single highest-probability grid point, which preserves discrete grid-scale features but may jump between adjacent cells. The mode-then-mean hybrid uses the mode to anchor the search region, then applies mean averaging for sub-grid precision.

How does the tracker handle cyclone formation without prior observations?

When no IBTrACS dataframe is provided, the tracker executes direct cyclogenesis by scanning prob_cyclone_exists for local maxima exceeding threshold, extracting candidates with _get_cyclogenesis_candidates, and pruning duplicates using min_disc_radius_between_cyclogenesis_candidates_km separation. This enables fully automated tracking on pure forecast outputs.

What determines when a track is terminated?

Tracks dissipate when the mean existence probability within the tracking disc falls below the user-specified threshold, checked in _advance_single_cyclone_active_at_lead_time_t (lines 48‑55). This probability-averaging approach is more robust than point-wise thresholds because it accounts for spatial uncertainty in the cyclone position.

Why does the tracker use Cartesian coordinates for averaging?

Latitude/longitude coordinates are non-Euclidean near the poles. The function _average_in_three_dimensions_and_project_on_sphere converts to 3D Cartesian coordinates via latlon_to_cartesian, computes the probability-weighted mean position in R³, then reprojects to the sphere surface with cartesian_to_latlon. This yields the true geodesic centroid rather than a naive arithmetic mean of latitude and longitude values.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →