# WeatherNext Computational Cost: TPU & GPU Pricing Breakdown for Google DeepMind's GenCast Model

> Discover the computational cost of Google DeepMind's WeatherNext model. Get a TPU and GPU pricing breakdown for GenCast and understand the cost of running forecasts.

- Repository: [Google DeepMind/weathernext](https://github.com/google-deepmind/weathernext)
- Tags: performance
- Published: 2026-08-16

---

**Running a single 30-step WeatherNext 0.25° forecast costs approximately $0.50–$2.22 on Google Cloud TPU v5p after one-time compilation, or $2.50–$11.10 for the initial compiled run.**

The computational cost of running WeatherNext depends primarily on hardware choice (TPU vs. GPU), model resolution (0.25° vs. 1°), and whether the model has been compiled. This guide breaks down the actual pricing, runtime estimates, and memory requirements based on the official Google DeepMind repository configuration.

## Key Cost Drivers for WeatherNext Inference

WeatherNext (specifically the **GenCast** forecast model) is optimized for Google Cloud TPU v5 devices, though GPU execution remains possible. Three factors dominate your compute bill:

- **Compilation phase** — one-time fixed cost per model version
- **Per-rollout inference time** — scales with forecast steps and chip count
- **Hardware hourly rate** — varies significantly between TPU v5e, v5p, and GPUs

## TPU Pricing and Runtime by Resolution

The [`docs/weathernext1_gen/cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/docs/weathernext1_gen/cloud_vm_setup.md) file contains authoritative pricing data. Current TPU rates (at time of repository publication) run approximately:

- **TPU v5e**: ~$1.20 per chip-hour
- **TPU v5p**: ~$5.20 per chip-hour

### 0.25° GenCast on TPU v5p (High Resolution)

This configuration delivers the most accurate forecasts but carries the highest compute cost.

| Metric | Value |
|--------|-------|
| Typical configuration | 8× v5p chips |
| Host RAM + HBM per chip | ~250 GB + 32 GB |
| Pre-compilation rollout time | ~30 minutes |
| Post-compilation rollout time | ~8 minutes |

**Cost per 30-step rollout:**
- Before compilation: **$2.50–$11.10**
- After compilation: **$0.50–$2.22**

The dramatic reduction occurs because compilation happens once per model version; subsequent rollouts reuse compiled kernels.

### 1° GenCast on TPU v5e (Lower Resolution)

This lighter configuration suits experimentation and faster iteration.

| Metric | Value |
|--------|-------|
| Typical configuration | 4× v5e chips |
| Host RAM + HBM per chip | ~21 GB + 8 GB |
| Pre-compilation rollout time | ~5 minutes |
| Post-compilation rollout time | Faster than pre-compilation |

**Cost per 30-step rollout:**
- Before compilation: **$0.11–$0.48**
- After compilation: **$0.07–$0.29**

## GPU Inference Costs

WeatherNext supports GPU execution via an alternative attention implementation (`triblockdiag_mha` in the source), but performance degrades significantly.

According to [`cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/cloud_vm_setup.md) (lines 94–97), GPU inference runs **approximately 2× slower** than equivalent TPU execution. A 30-step 0.25° rollout requiring ~8 minutes on TPU v5p takes ~25 minutes on comparable H100 hardware.

Memory requirements also increase substantially for GPU execution:
- Host RAM: ~300 GB
- HBM: ~60 GB

No explicit per-hour GPU pricing is documented in the repository, but you can estimate costs by applying your cloud provider's H100 rate to the doubled runtime. For example, at $3/hour for an H100, a 25-minute rollout costs roughly $1.25 in GPU time versus $0.69–$2.22 for the faster TPU v5p run.

## Understanding Compilation vs. Runtime Costs

The **one-time compilation cost** represents a critical optimization opportunity. As documented in [`cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/cloud_vm_setup.md) (lines 27–35):

> The compilation/tracing phase incurs a fixed cost that does *not* increase with the number of devices.

This means:
- First run of a new model version: pay full compilation + runtime
- All subsequent runs: pay reduced runtime only

For production forecasting pipelines, the compilation cost amortizes quickly across hundreds or thousands of rollouts.

## Estimating WeatherNext Costs Programmatically

The repository includes utilities for cost projection. Below are working examples based on [`weathernext/utils/model_utils.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/utils/model_utils.py):

```python

# Estimate pre-compilation cost for 0.25° GenCast on v5p-8

chip_price_per_hour = 5.20          # $/chip hour (v5p)

rollout_minutes = 30                # wall-clock time before compilation

chips = 8

cost = (rollout_minutes / 60) * chip_price_per_hour * chips
print(f"Pre-compilation rollout cost: ${cost:.2f}")

# Output: Pre-compilation rollout cost: $20.80

```

```python

# Same configuration after compilation (8-minute runtime)

rollout_minutes = 8
cost = (rollout_minutes / 60) * chip_price_per_hour * chips
print(f"Post-compilation rollout cost: ${cost:.2f}")

# Output: Post-compilation rollout cost: $5.57

```

Note: The repository's documented cost ranges ($0.50–$2.22, $2.50–$11.10) assume idealized conditions and potential batching efficiencies; your actual costs may vary based on cloud region, sustained use discounts, and specific pod configuration.

## Memory Requirements Summary

| Resolution | Hardware | Host RAM | Per-Chip HBM |
|------------|----------|----------|--------------|
| 0.25° GenCast | TPU v5p | ~250 GB | 32 GB |
| 1° GenCast | TPU v5e | ~21 GB | 8 GB |
| 0.25° (GPU) | H100 | ~300 GB | 60 GB |

These figures from [`cloud_vm_setup.md`](https://github.com/google-deepmind/weathernext/blob/main/cloud_vm_setup.md) (lines 44–46, 88–92) determine your minimum instance sizing regardless of cost optimization strategy.

## Summary

- **WeatherNext computational cost** scales with TPU chip count, hourly rate, and rollout duration
- **TPU v5p-8 configuration** for 0.25° GenCast: $0.50–$2.22 per rollout after compilation
- **TPU v5e-4 configuration** for 1° GenCast: $0.07–$0.29 per rollout after compilation
- **Compilation** is a one-time cost that dramatically reduces per-run pricing
- **GPU inference** costs roughly double in wall-clock time versus equivalent TPU runs

For production deployments, prioritize TPU execution and plan to amortize compilation costs across many forecast generations.

## Frequently Asked Questions

### What is the cheapest way to run WeatherNext?

The **1° GenCast model on TPU v5e** offers the lowest per-rollout cost at approximately $0.07–$0.29 after compilation, according to the repository's cloud setup documentation. This configuration requires only 4 chips and minimal memory compared to the high-resolution alternative.

### How long does WeatherNext compilation take?

Compilation time varies by resolution and hardware, but the **cost is fixed regardless of device count**. For 0.25° GenCast, expect roughly 30 minutes of wall-clock time before the model is ready for optimized inference; for 1° GenCast, approximately 5 minutes. This one-time investment enables 60–80% cost reductions on subsequent rollouts.

### Can I run WeatherNext on consumer GPUs?

The repository documents **H100-level GPU execution** as a fallback option, requiring ~300 GB host RAM and ~60 GB HBM. Consumer GPUs lack sufficient memory for the 0.25° model, and the `triblockdiag_mha` attention implementation performs substantially slower than TPU-native code. Cloud TPUs remain the recommended production path.

### What determines WeatherNext inference speed?

Three factors control rollout duration in [`weathernext/weathernext1_gen/gencast.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext1_gen/gencast.py): the **number of forecast steps** (default 30), **hardware type** (TPU v5p vs. v5e vs. GPU), and **whether kernels are compiled**. The underlying GraphCast architecture in [`weathernext/weathernext1_graph/graphcast.py`](https://github.com/google-deepmind/weathernext/blob/main/weathernext/weathernext1_graph/graphcast.py) defines the computational graph that executes at each step.