# How to Deploy Marin to Production: A Complete Guide to Pulumi-Based Rollouts

> Deploy Marin to production using Pulumi-based rollouts with the marin-deploy CLI or GitHub Actions. Follow this guide for a seamless service deployment.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-27

---

**To deploy Marin to production, use the `marin-deploy` CLI operator to run Pulumi-based rollouts for individual services (ducky, echo, finelog, etc.) after syncing your local environment with `uv sync --all-packages --extra deploy`, or rely on the automated GitHub Actions workflow that triggers on merges to `main`.**

Marin is an open-source machine learning infrastructure platform that orchestrates workloads across pre-emptible GCP TPUs. Deploying Marin to production involves managing infrastructure-as-code through Pulumi stacks located in the `infra/` directory of the `marin-community/marin` repository.

## Production Architecture Overview

Marin runs on **pre-emptible GCP TPU clusters** managed by two core orchestration systems: **Iris** for job scheduling and cluster provisioning, and **Zephyr** for data pipeline abstractions. The production environment spans multiple clusters including `marin-us-central2` and `marin-us-west4`, as documented in [`infra/README.md`](https://github.com/marin-community/marin/blob/main/infra/README.md).

Each application service—such as `ducky`, `echo`, `finelog`, `grafana`, `loom`, and `xprof`—is defined as an independent Pulumi project under `infra/<service>/`. This modular architecture allows you to deploy services independently without affecting the broader cluster infrastructure.

## Prepare Your Local Environment

Before deploying to production, ensure your local environment matches the repository's toolchain requirements. Marin uses `uv` for Python package management and requires the `deploy` extra group for Pulumi operations.

Install all dependencies and the deployment toolchain:

```bash
uv sync --all-packages --extra deploy

```

Verify that the **Pulumi CLI** is available in your path, as `marin-deploy` invokes Pulumi programmatically to apply infrastructure changes.

## Deploy Services with the marin-deploy Operator

The `marin-deploy` command serves as the primary interface for production rollouts. Located in the infrastructure deployment system (documented in [`infra/deploy/README.md`](https://github.com/marin-community/marin/blob/main/infra/deploy/README.md)), this operator wraps Pulumi commands to handle stack selection, configuration overrides, and secret resolution automatically.

### Basic Rollout Command

To deploy or redeploy a specific service to production:

```bash
uv run --all-packages --extra deploy marin-deploy ducky rollout

```

This command performs a **preview** of changes followed by an **apply** operation against the production stack for the `ducky` service. The operator ensures that the deployed GCP resources—including GKE clusters, GCS buckets, and IAM bindings—match the declarations in `infra/ducky/`.

### Deploying with Configuration Overrides

You can pass runtime configuration values using the `--config` flag. These overrides apply to a temporary copy of the stack configuration, leaving the checked-in `Pulumi.<stack>.yaml` untouched:

```bash
uv run --all-packages --extra deploy \
  marin-deploy ducky rollout --config enable_new_feature=true

```

### Non-Interactive Deployment

For CI/CD pipelines or automated scripts, skip Pulumi's confirmation prompt with the `--yes` flag:

```bash
uv run --all-packages --extra deploy marin-deploy echo rollout --yes

```

## Secrets and Configuration Management

Production secrets—such as Cloudflare DNS tokens and API credentials—are **never stored in the repository**. Instead, `marin-deploy` fetches service-specific secrets from **GCP Secret Manager** at rollout time. This approach ensures that sensitive values remain encrypted and access-controlled outside of version control.

Configuration overrides specified via `--config KEY=VALUE` are applied transiently during the deployment process, allowing you to modify feature flags or resource allocations without committing changes to the infrastructure code.

## Automated CI/CD Deployment

While manual rollouts provide granular control, Marin also supports fully automated deployments through GitHub Actions. The workflow defined in [`.github/workflows/ops-pulumi-rollout.yaml`](https://github.com/marin-community/marin/blob/main/.github/workflows/ops-pulumi-rollout.yaml) monitors the `main` branch and automatically triggers production rollouts when pull requests are merged.

This CI pipeline runs the same `marin-deploy` commands used locally but executes with production credentials stored as repository secrets. The automated approach ensures that the deployed infrastructure always reflects the current state of the `main` branch.

## Rollback Procedures

If a deployment introduces instability, each service supports a `rollback` sub-command that leverages Pulumi's state history to revert to previous revisions.

### Rollback to Previous Revision

To revert the most recent deployment of a service:

```bash
uv run --all-packages --extra deploy marin-deploy finelog rollback

```

### Rollback to Specific Revision

To target a specific historical revision:

```bash
uv run --all-packages --extra deploy \
  marin-deploy finelog rollback --to-revision 12

```

The rollback process restores both the previous infrastructure state and any checkpointed job data, which is critical for maintaining consistency in Marin’s pre-emptible TPU environment where jobs must handle interruption gracefully.

## Summary

Deploying Marin to production relies on a modular, Pulumi-based infrastructure system managed through the `marin-deploy` operator. Key points to remember:

- **Install dependencies** with `uv sync --all-packages --extra deploy` before running any deployment commands
- **Deploy individual services** using `marin-deploy <service> rollout`, which automatically handles Pulumi previews and applies
- **Use configuration overrides** with `--config` for temporary changes without modifying checked-in YAML files
- **Fetch secrets from GCP Secret Manager** at deployment time rather than storing them in the repository
- **Enable automated rollouts** via the GitHub Actions workflow in [`.github/workflows/ops-pulumi-rollout.yaml`](https://github.com/marin-community/marin/blob/main/.github/workflows/ops-pulumi-rollout.yaml) for continuous deployment
- **Execute rollbacks** with the `rollback` sub-command, optionally specifying a target revision with `--to-revision`

## Frequently Asked Questions

### What is the difference between marin-deploy and the standard Pulumi CLI?

`marin-deploy` is a specialized operator that wraps the Pulumi CLI with Marin-specific logic. According to the source code in [`infra/deploy/README.md`](https://github.com/marin-community/marin/blob/main/infra/deploy/README.md), it handles stack selection for production environments, automatically resolves secrets from GCP Secret Manager, and applies configuration overrides to temporary stack copies. While you could use `pulumi up` directly, `marin-deploy` ensures consistent handling of Marin’s service dependencies and cluster contexts.

### Can I deploy multiple services simultaneously?

The `marin-deploy` operator is designed to target **single services** (e.g., `ducky`, `echo`) for independent rollouts. To deploy multiple services, you must invoke the command separately for each service. This intentional limitation prevents cascading failures across the production TPU clusters and allows staged rollouts where you verify each service before proceeding to the next.

### Where are production secrets stored in the Marin repository?

Production secrets are **not stored in the repository**. As implemented in the deployment operator, service-specific credentials like Cloudflare DNS tokens are fetched dynamically from **GCP Secret Manager** during the rollout process. The infrastructure code in `infra/` only references secret names, never plaintext values, ensuring that sensitive data remains encrypted and access-controlled through Google Cloud IAM policies.

### How does Marin handle deployments to pre-emptible TPU infrastructure?

Marin’s production environment runs primarily on pre-emptible GCP TPUs, which can be interrupted at any time. The deployment process documented in [`infra/README.md`](https://github.com/marin-community/marin/blob/main/infra/README.md) emphasizes **fast start-up** and **idempotent, checkpoint-able jobs**. The `marin-deploy` operator ensures that rolled-out services can recover quickly from preemption events, and the rollback functionality preserves checkpointed states to minimize data loss during infrastructure reversions.