WeatherNext Computational Cost: TPU & GPU Pricing Breakdown for Google DeepMind's GenCast Model
Running a single 30-step WeatherNext 0.25° forecast costs approximately $0.50–$2.22 on Google Cloud TPU v5p after one-time compilation, or $2.50–$11.10 for the initial compiled run.
The computational cost of running WeatherNext depends primarily on hardware choice (TPU vs. GPU), model resolution (0.25° vs. 1°), and whether the model has been compiled. This guide breaks down the actual pricing, runtime estimates, and memory requirements based on the official Google DeepMind repository configuration.
Key Cost Drivers for WeatherNext Inference
WeatherNext (specifically the GenCast forecast model) is optimized for Google Cloud TPU v5 devices, though GPU execution remains possible. Three factors dominate your compute bill:
- Compilation phase — one-time fixed cost per model version
- Per-rollout inference time — scales with forecast steps and chip count
- Hardware hourly rate — varies significantly between TPU v5e, v5p, and GPUs
TPU Pricing and Runtime by Resolution
The docs/weathernext1_gen/cloud_vm_setup.md file contains authoritative pricing data. Current TPU rates (at time of repository publication) run approximately:
- TPU v5e: ~$1.20 per chip-hour
- TPU v5p: ~$5.20 per chip-hour
0.25° GenCast on TPU v5p (High Resolution)
This configuration delivers the most accurate forecasts but carries the highest compute cost.
| Metric | Value |
|---|---|
| Typical configuration | 8× v5p chips |
| Host RAM + HBM per chip | ~250 GB + 32 GB |
| Pre-compilation rollout time | ~30 minutes |
| Post-compilation rollout time | ~8 minutes |
Cost per 30-step rollout:
- Before compilation: $2.50–$11.10
- After compilation: $0.50–$2.22
The dramatic reduction occurs because compilation happens once per model version; subsequent rollouts reuse compiled kernels.
1° GenCast on TPU v5e (Lower Resolution)
This lighter configuration suits experimentation and faster iteration.
| Metric | Value |
|---|---|
| Typical configuration | 4× v5e chips |
| Host RAM + HBM per chip | ~21 GB + 8 GB |
| Pre-compilation rollout time | ~5 minutes |
| Post-compilation rollout time | Faster than pre-compilation |
Cost per 30-step rollout:
- Before compilation: $0.11–$0.48
- After compilation: $0.07–$0.29
GPU Inference Costs
WeatherNext supports GPU execution via an alternative attention implementation (triblockdiag_mha in the source), but performance degrades significantly.
According to cloud_vm_setup.md (lines 94–97), GPU inference runs approximately 2× slower than equivalent TPU execution. A 30-step 0.25° rollout requiring ~8 minutes on TPU v5p takes ~25 minutes on comparable H100 hardware.
Memory requirements also increase substantially for GPU execution:
- Host RAM: ~300 GB
- HBM: ~60 GB
No explicit per-hour GPU pricing is documented in the repository, but you can estimate costs by applying your cloud provider's H100 rate to the doubled runtime. For example, at $3/hour for an H100, a 25-minute rollout costs roughly $1.25 in GPU time versus $0.69–$2.22 for the faster TPU v5p run.
Understanding Compilation vs. Runtime Costs
The one-time compilation cost represents a critical optimization opportunity. As documented in cloud_vm_setup.md (lines 27–35):
The compilation/tracing phase incurs a fixed cost that does not increase with the number of devices.
This means:
- First run of a new model version: pay full compilation + runtime
- All subsequent runs: pay reduced runtime only
For production forecasting pipelines, the compilation cost amortizes quickly across hundreds or thousands of rollouts.
Estimating WeatherNext Costs Programmatically
The repository includes utilities for cost projection. Below are working examples based on weathernext/utils/model_utils.py:
# Estimate pre-compilation cost for 0.25° GenCast on v5p-8
chip_price_per_hour = 5.20 # $/chip hour (v5p)
rollout_minutes = 30 # wall-clock time before compilation
chips = 8
cost = (rollout_minutes / 60) * chip_price_per_hour * chips
print(f"Pre-compilation rollout cost: ${cost:.2f}")
# Output: Pre-compilation rollout cost: $20.80
# Same configuration after compilation (8-minute runtime)
rollout_minutes = 8
cost = (rollout_minutes / 60) * chip_price_per_hour * chips
print(f"Post-compilation rollout cost: ${cost:.2f}")
# Output: Post-compilation rollout cost: $5.57
Note: The repository's documented cost ranges ($0.50–$2.22, $2.50–$11.10) assume idealized conditions and potential batching efficiencies; your actual costs may vary based on cloud region, sustained use discounts, and specific pod configuration.
Memory Requirements Summary
| Resolution | Hardware | Host RAM | Per-Chip HBM |
|---|---|---|---|
| 0.25° GenCast | TPU v5p | ~250 GB | 32 GB |
| 1° GenCast | TPU v5e | ~21 GB | 8 GB |
| 0.25° (GPU) | H100 | ~300 GB | 60 GB |
These figures from cloud_vm_setup.md (lines 44–46, 88–92) determine your minimum instance sizing regardless of cost optimization strategy.
Summary
- WeatherNext computational cost scales with TPU chip count, hourly rate, and rollout duration
- TPU v5p-8 configuration for 0.25° GenCast: $0.50–$2.22 per rollout after compilation
- TPU v5e-4 configuration for 1° GenCast: $0.07–$0.29 per rollout after compilation
- Compilation is a one-time cost that dramatically reduces per-run pricing
- GPU inference costs roughly double in wall-clock time versus equivalent TPU runs
For production deployments, prioritize TPU execution and plan to amortize compilation costs across many forecast generations.
Frequently Asked Questions
What is the cheapest way to run WeatherNext?
The 1° GenCast model on TPU v5e offers the lowest per-rollout cost at approximately $0.07–$0.29 after compilation, according to the repository's cloud setup documentation. This configuration requires only 4 chips and minimal memory compared to the high-resolution alternative.
How long does WeatherNext compilation take?
Compilation time varies by resolution and hardware, but the cost is fixed regardless of device count. For 0.25° GenCast, expect roughly 30 minutes of wall-clock time before the model is ready for optimized inference; for 1° GenCast, approximately 5 minutes. This one-time investment enables 60–80% cost reductions on subsequent rollouts.
Can I run WeatherNext on consumer GPUs?
The repository documents H100-level GPU execution as a fallback option, requiring ~300 GB host RAM and ~60 GB HBM. Consumer GPUs lack sufficient memory for the 0.25° model, and the triblockdiag_mha attention implementation performs substantially slower than TPU-native code. Cloud TPUs remain the recommended production path.
What determines WeatherNext inference speed?
Three factors control rollout duration in weathernext/weathernext1_gen/gencast.py: the number of forecast steps (default 30), hardware type (TPU v5p vs. v5e vs. GPU), and whether kernels are compiled. The underlying GraphCast architecture in weathernext/weathernext1_graph/graphcast.py defines the computational graph that executes at each step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →