How to Deploy Marin to Production: A Complete Guide to Pulumi-Based Rollouts
To deploy Marin to production, use the marin-deploy CLI operator to run Pulumi-based rollouts for individual services (ducky, echo, finelog, etc.) after syncing your local environment with uv sync --all-packages --extra deploy, or rely on the automated GitHub Actions workflow that triggers on merges to main.
Marin is an open-source machine learning infrastructure platform that orchestrates workloads across pre-emptible GCP TPUs. Deploying Marin to production involves managing infrastructure-as-code through Pulumi stacks located in the infra/ directory of the marin-community/marin repository.
Production Architecture Overview
Marin runs on pre-emptible GCP TPU clusters managed by two core orchestration systems: Iris for job scheduling and cluster provisioning, and Zephyr for data pipeline abstractions. The production environment spans multiple clusters including marin-us-central2 and marin-us-west4, as documented in infra/README.md.
Each application service—such as ducky, echo, finelog, grafana, loom, and xprof—is defined as an independent Pulumi project under infra/<service>/. This modular architecture allows you to deploy services independently without affecting the broader cluster infrastructure.
Prepare Your Local Environment
Before deploying to production, ensure your local environment matches the repository's toolchain requirements. Marin uses uv for Python package management and requires the deploy extra group for Pulumi operations.
Install all dependencies and the deployment toolchain:
uv sync --all-packages --extra deploy
Verify that the Pulumi CLI is available in your path, as marin-deploy invokes Pulumi programmatically to apply infrastructure changes.
Deploy Services with the marin-deploy Operator
The marin-deploy command serves as the primary interface for production rollouts. Located in the infrastructure deployment system (documented in infra/deploy/README.md), this operator wraps Pulumi commands to handle stack selection, configuration overrides, and secret resolution automatically.
Basic Rollout Command
To deploy or redeploy a specific service to production:
uv run --all-packages --extra deploy marin-deploy ducky rollout
This command performs a preview of changes followed by an apply operation against the production stack for the ducky service. The operator ensures that the deployed GCP resources—including GKE clusters, GCS buckets, and IAM bindings—match the declarations in infra/ducky/.
Deploying with Configuration Overrides
You can pass runtime configuration values using the --config flag. These overrides apply to a temporary copy of the stack configuration, leaving the checked-in Pulumi.<stack>.yaml untouched:
uv run --all-packages --extra deploy \
marin-deploy ducky rollout --config enable_new_feature=true
Non-Interactive Deployment
For CI/CD pipelines or automated scripts, skip Pulumi's confirmation prompt with the --yes flag:
uv run --all-packages --extra deploy marin-deploy echo rollout --yes
Secrets and Configuration Management
Production secrets—such as Cloudflare DNS tokens and API credentials—are never stored in the repository. Instead, marin-deploy fetches service-specific secrets from GCP Secret Manager at rollout time. This approach ensures that sensitive values remain encrypted and access-controlled outside of version control.
Configuration overrides specified via --config KEY=VALUE are applied transiently during the deployment process, allowing you to modify feature flags or resource allocations without committing changes to the infrastructure code.
Automated CI/CD Deployment
While manual rollouts provide granular control, Marin also supports fully automated deployments through GitHub Actions. The workflow defined in .github/workflows/ops-pulumi-rollout.yaml monitors the main branch and automatically triggers production rollouts when pull requests are merged.
This CI pipeline runs the same marin-deploy commands used locally but executes with production credentials stored as repository secrets. The automated approach ensures that the deployed infrastructure always reflects the current state of the main branch.
Rollback Procedures
If a deployment introduces instability, each service supports a rollback sub-command that leverages Pulumi's state history to revert to previous revisions.
Rollback to Previous Revision
To revert the most recent deployment of a service:
uv run --all-packages --extra deploy marin-deploy finelog rollback
Rollback to Specific Revision
To target a specific historical revision:
uv run --all-packages --extra deploy \
marin-deploy finelog rollback --to-revision 12
The rollback process restores both the previous infrastructure state and any checkpointed job data, which is critical for maintaining consistency in Marin’s pre-emptible TPU environment where jobs must handle interruption gracefully.
Summary
Deploying Marin to production relies on a modular, Pulumi-based infrastructure system managed through the marin-deploy operator. Key points to remember:
- Install dependencies with
uv sync --all-packages --extra deploybefore running any deployment commands - Deploy individual services using
marin-deploy <service> rollout, which automatically handles Pulumi previews and applies - Use configuration overrides with
--configfor temporary changes without modifying checked-in YAML files - Fetch secrets from GCP Secret Manager at deployment time rather than storing them in the repository
- Enable automated rollouts via the GitHub Actions workflow in
.github/workflows/ops-pulumi-rollout.yamlfor continuous deployment - Execute rollbacks with the
rollbacksub-command, optionally specifying a target revision with--to-revision
Frequently Asked Questions
What is the difference between marin-deploy and the standard Pulumi CLI?
marin-deploy is a specialized operator that wraps the Pulumi CLI with Marin-specific logic. According to the source code in infra/deploy/README.md, it handles stack selection for production environments, automatically resolves secrets from GCP Secret Manager, and applies configuration overrides to temporary stack copies. While you could use pulumi up directly, marin-deploy ensures consistent handling of Marin’s service dependencies and cluster contexts.
Can I deploy multiple services simultaneously?
The marin-deploy operator is designed to target single services (e.g., ducky, echo) for independent rollouts. To deploy multiple services, you must invoke the command separately for each service. This intentional limitation prevents cascading failures across the production TPU clusters and allows staged rollouts where you verify each service before proceeding to the next.
Where are production secrets stored in the Marin repository?
Production secrets are not stored in the repository. As implemented in the deployment operator, service-specific credentials like Cloudflare DNS tokens are fetched dynamically from GCP Secret Manager during the rollout process. The infrastructure code in infra/ only references secret names, never plaintext values, ensuring that sensitive data remains encrypted and access-controlled through Google Cloud IAM policies.
How does Marin handle deployments to pre-emptible TPU infrastructure?
Marin’s production environment runs primarily on pre-emptible GCP TPUs, which can be interrupted at any time. The deployment process documented in infra/README.md emphasizes fast start-up and idempotent, checkpoint-able jobs. The marin-deploy operator ensures that rolled-out services can recover quickly from preemption events, and the rollback functionality preserves checkpointed states to minimize data loss during infrastructure reversions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →