Direct Cyclone Tracker Implementation and Track Extraction in WeatherNext
The DirectTracker class in weathernext/cyclones/direct_tracker.py implements a greedy cyclone-tracking algorithm that initializes cyclone centers via IBTrACS data or direct cyclogenesis, evolves centers through time using momentum-based prediction and probability-weighted refinement, and returns tracks as pandas DataFrames.
The direct cyclone tracker is a core component of the WeatherNext ensemble forecasting system developed by Google DeepMind. It transforms raw gridded probabilistic forecasts into discrete tropical cyclone trajectories without requiring external reanalysis data. This article explains how the tracker is architected and how to extract complete storm tracks from model outputs.
Direct Tracker Architecture
The DirectTracker follows a three-stage pipeline: initialization, temporal evolution, and result collection. Each stage uses geodesic distance calculations and probability-weighted operations on the sphere.
Stage 1: Cyclone Center Initialization
The tracker accepts two initialization modes controlled by the __init__ method (lines 33‑59).
- IBTrACS-guided mode: When a historical dataframe is supplied, centers are seeded at observed locations
- Direct cyclogenesis: The algorithm scans
prob_cyclone_existsfields for high-probability blobs, ranks candidates, and prunes them using minimum separation distances
The cyclogenesis logic resides in _do_direct_cyclogenesis (lines 78‑74), which delegates candidate detection to _get_cyclogenesis_candidates (lines 70‑44) and refinement to _prune_and_refine_cyclogenesis_candidates (lines 45‑92). The refinement step uses disc_radius_cyclogenesis_refinement_km to merge nearby candidates and select probability maxima.
Stage 2: Temporal Evolution with Momentum
For each lead-time step, active cyclones are advanced through four operations in _advance_single_cyclone_active_at_lead_time_t (lines 109‑124):
-
Momentum prediction:
_get_latlon_guess_with_momentum_update(lines 76‑99) extrapolates the next position using the current velocity vector scaled bymomentum_constant -
Tracking mode refinement: The predicted position is corrected using one of three modes operating on a disc-shaped neighborhood:
- mean: Probability-weighted average via
_get_mean_latlon_of_grid_points_within_disc - mode: Highest-probability grid point via
_get_mode_latlon_of_grid_points_within_disc - mode-then-mean: Mode seeding followed by mean refinement
- mean: Probability-weighted average via
-
Scalar variable extraction: Two methods extract intensity variables:
_get_scalar_variables_via_bilinear_interpolation(lines 96‑102)_get_all_scalar_variables_within_discfor probability-weighted disc averages
-
Dissipation check: Tracks terminate when mean existence probability falls below threshold (lines 48‑55)
Stage 3: Track Assembly
Each center update generates a row via _single_row_dataframe_from_dict (lines 83‑85). Rows accumulate in lists and concatenate through pd.concat into a complete trajectory DataFrame containing latitude, longitude, lead time, valid time, wind speed, pressure, radii, and existence probability.
Core Geospatial Operations
The tracker performs all spatial calculations on the sphere using these helpers:
| Function | Purpose | Location |
|---|---|---|
_get_bounding_box_sides_in_degrees |
Computes lat/lon box containing a geodesic disc of radius r km | direct_tracker.py lines 87‑34 |
_get_is_in_radius_grid_around_latlon |
Boolean mask for points inside a disc around a center | direct_tracker.py lines 47‑33 |
_average_in_three_dimensions_and_project_on_sphere |
Probability-weighted averaging on unit sphere | direct_tracker.py lines 37‑22 |
slice_data_array_latlon_box_with_lon_wraparound |
Sub-grid extraction handling 0-360° wrap-around | utils/tracker_utils.py |
latlon_to_cartesian / cartesian_to_latlon |
Spherical coordinate conversions | utils/cyclone_utils.py |
Extracting Cyclone Tracks: Complete Workflow
Step 1: Load Configuration
from weathernext.cyclones.direct_tracker_6h_v1_config import get_config
cfg = get_config() # Pre-tuned for 6-hour forecast cadence
The configuration object bundles all radii, thresholds, and the tracking mode selection.
Step 2: Instantiate Tracker
tracker = cfg.tracker_constructor(
disc_radius_scalar_variables_km=cfg.disc_radius_scalar_variables_km,
disc_radius_mean_latlon_km=cfg.disc_radius_mean_latlon_km,
disc_radius_mode_latlon_km=cfg.disc_radius_mode_latlon_km,
disc_radius_mean_probability_of_existence_km=cfg.disc_radius_mean_probability_of_existence_km,
disc_radius_cyclogenesis_refinement_km=cfg.disc_radius_cyclogenesis_refinement_km,
tracking_mode=cfg.tracking_mode, # 'mean', 'mode', or 'mode_then_mean'
momentum_constant=cfg.momentum_constant,
min_disc_radius_between_cyclogenesis_candidates_km=cfg.min_disc_radius_between_cyclogenesis_candidates_km,
temporal_resolution_hours=cfg.temporal_resolution_hours,
)
Step 3: Run Tracking on Gridded Forecast
track_df = tracker.track(gridded_ds) # gridded_ds: xr.Dataset with cyclone variables
The track method is defined in the abstract base class CycloneTracker (tracker_base.py) and orchestrates the three-stage pipeline.
Step 4: Analyze Output
print(track_df.columns)
# Output: ['lat', 'lon', 'lead_time', 'valid_time', 'max_sustained_wind_speed_knots',
# 'radius_of_maximum_winds', 'prob_cyclone_exists', ...]
print(track_df[['lat', 'lon', 'max_sustained_wind_speed_knots']].head())
Key Source Files
| File | Responsibility |
|---|---|
weathernext/cyclones/direct_tracker.py |
Main DirectTracker implementation with cyclogenesis, momentum updates, and tracking modes |
weathernext/cyclones/direct_tracker_6h_v1_config.py |
Production configuration for 6-hourly forecasts |
weathernext/cyclones/tracker_base.py |
Abstract CycloneTracker class defining the track API |
weathernext/cyclones/tracker_utils.py |
Dataset slicing and interpolation utilities |
weathernext/cyclones/cyclone_utils.py |
Geodesic distances and spherical coordinate math |
weathernext/cyclones/ibtracs_processing_utils.py |
IBTrACS to gridded variable name mapping |
weathernext/cyclones/constants.py |
Column name and coordinate key definitions |
Summary
- The direct cyclone tracker implements greedy trajectory extraction through probability-weighted operations on forecast grids
- Three tracking modes (mean, mode, mode-then-mean) control how center positions are refined from gridded probabilities
- Momentum-based prediction reduces search windows and maintains track continuity across lead times
- Tracks extract as pandas DataFrames with full temporal evolution of position, intensity, and existence probability
- The 6-hour configuration provides production-ready parameters without manual tuning
Frequently Asked Questions
What is the difference between mean and mode tracking modes?
Mean tracking computes the probability-weighted average of all grid points within the search disc using _average_in_three_dimensions_and_project_on_sphere, producing smooth trajectories. Mode tracking selects the single highest-probability grid point, which preserves discrete grid-scale features but may jump between adjacent cells. The mode-then-mean hybrid uses the mode to anchor the search region, then applies mean averaging for sub-grid precision.
How does the tracker handle cyclone formation without prior observations?
When no IBTrACS dataframe is provided, the tracker executes direct cyclogenesis by scanning prob_cyclone_exists for local maxima exceeding threshold, extracting candidates with _get_cyclogenesis_candidates, and pruning duplicates using min_disc_radius_between_cyclogenesis_candidates_km separation. This enables fully automated tracking on pure forecast outputs.
What determines when a track is terminated?
Tracks dissipate when the mean existence probability within the tracking disc falls below the user-specified threshold, checked in _advance_single_cyclone_active_at_lead_time_t (lines 48‑55). This probability-averaging approach is more robust than point-wise thresholds because it accounts for spatial uncertainty in the cyclone position.
Why does the tracker use Cartesian coordinates for averaging?
Latitude/longitude coordinates are non-Euclidean near the poles. The function _average_in_three_dimensions_and_project_on_sphere converts to 3D Cartesian coordinates via latlon_to_cartesian, computes the probability-weighted mean position in R³, then reprojects to the sphere surface with cartesian_to_latlon. This yields the true geodesic centroid rather than a naive arithmetic mean of latitude and longitude values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →