How to Deploy Models with Marin-Deploy: A Complete Guide to ML Model Serving
Marin-Deploy is the command-line entry point that ships with the marin-deploy package, providing a unified interface for deploying machine-learning models to Kubernetes or cloud platforms using declarative YAML configurations.
Deploying machine-learning models to production requires reproducible infrastructure and streamlined workflows. Marin-Deploy, part of the marin-community/marin repository, is the official CLI tool designed specifically to deploy models with Marin-Deploy configurations across various backends including Kubernetes and Google Cloud Run.
Understanding Marin-Deploy Architecture
Marin-Deploy operates as a Click-based CLI defined in infra/deploy/src/marin_deploy/cli.py. The tool registers multiple sub-commands that handle different deployment targets. When you deploy models with Marin-Deploy, you interact with two primary components:
- Finelog: The primary model-serving deployment system implemented in
infra/deploy/src/marin_deploy/finelog.py - Pulumi Services: Cloud-specific abstractions defined in
infra/deploy/src/marin_deploy/pulumi.pyfor CoreWeave, GCP Cloud Run, and other platforms
Installation and Setup
Before you can deploy models with Marin-Deploy, install the package from the infrastructure directory:
uv pip install -e infra/deploy
This installs the marin-deploy command-line entry point and its dependencies, including the Finelog configuration loader and Pulumi integration libraries.
Creating a Finelog Deployment Configuration
Model deployments in Marin rely on YAML configuration files. The configuration schema, defined in lib/finelog/src/finelog/deploy/config.py (lines 65-227), specifies container images, resource requirements, and Kubernetes namespaces.
Configuration Schema Structure
A valid Finelog configuration requires the following structure:
deployment:
name: my-model-service
k8s:
namespace: models
image: gcr.io/my-project/my-model:latest
resources:
cpu: "4"
memory: "16Gi"
The load_finelog_config function—referenced in utilities like scripts/ops/storage/coreweave_usage.py—validates this YAML against the schema and loads it into memory for processing.
Deploying Models with the Finelog Command
The finelog sub-command is the primary method to deploy models with Marin-Deploy to Kubernetes clusters.
Basic Model Deployment
Execute the deployment using your configuration file:
marin-deploy finelog --config finelog/config/my_model.yaml
When invoked, the CLI performs three critical operations:
- Loads the YAML configuration using
load_finelog_config - Resolves the container image to an immutable digest using the helper in
lib/finelog/src/finelog/deploy/image.py, ensuring reproducible deployments - Creates or updates the Kubernetes Deployment through Pulumi abstractions defined in
infra/deploy/src/marin_deploy/pulumi.py
Validating Deployments with Dry-Run
To inspect what would be deployed without applying changes:
marin-deploy finelog --config finelog/config/my_model.yaml --dry-run
This validates your configuration and previews resource changes without modifying cluster state.
Rollback and Revision Management
Marin-Deploy tracks deployment revisions using Kubernetes annotations. The revision-tracking logic resides in lib/finelog/src/finelog/deploy/_k8s.py (lines 187-263).
Performing Rollbacks
To rollback to a previous revision:
# Identify the target revision number
kubectl get deployment sentiment-analyzer -n ml-models -o=jsonpath='{.metadata.annotations.finlog\.revision}'
# Execute rollback
marin-deploy finelog --config finelog/config/my_model.yaml --rollback 3
The workflow is deliberately idempotent—running the same command twice reconciles the live state with your YAML configuration without recreating resources unnecessarily.
Alternative Deployment Targets
While Finelog targets Kubernetes, Marin-Deploy supports Pulumi-backed cloud services through the service group registered in infra/deploy/src/marin_deploy/pulumi.py.
Deploying to Google Cloud Run
To deploy the same model configuration to GCP Cloud Run instead of Kubernetes:
marin-deploy cloud-run --config finelog/config/my_model.yaml
This leverages the same YAML configuration but provisions serverless infrastructure through Pulumi rather than Kubernetes Deployments.
Summary
- Marin-Deploy provides a unified CLI for deploying ML models through the
marin-deploypackage - Deployments rely on Finelog YAML configurations specifying container images, resources, and namespaces
- The
finelogsub-command handles Kubernetes deployments, while Pulumi services manage cloud platforms like Cloud Run - Image resolution pins container tags to immutable digests for reproducibility
- Idempotent operations ensure safe repeated executions without unnecessary resource recreation
- Rollback capabilities leverage Kubernetes annotations tracked in
lib/finelog/src/finelog/deploy/_k8s.py
Frequently Asked Questions
What is the difference between marin-deploy finelog and marin-deploy cloud-run?
The finelog sub-command deploys models to Kubernetes clusters using native Deployment resources, while cloud-run and other Pulumi-backed services provision serverless infrastructure on specific cloud providers. Both use the same YAML configuration format defined in lib/finelog/src/finelog/deploy/config.py, but target different execution environments through the abstractions in infra/deploy/src/marin_deploy/pulumi.py.
How does Marin-Deploy ensure reproducible model deployments?
Marin-Deploy enforces reproducibility through immutable image resolution. The helper in lib/finelog/src/finelog/deploy/image.py resolves container tags to specific digests before deployment, ensuring that gcr.io/project/model:latest translates to an immutable SHA256 reference. This prevents accidental deployment of untested image versions.
Can I deploy multiple models using a single configuration file?
While each YAML configuration typically defines a single deployment resource, you can orchestrate multiple model services by creating separate configuration files under lib/finelog/config/ and invoking marin-deploy finelog separately for each. The tool processes one deployment configuration per invocation to maintain clear resource boundaries and rollback capabilities.
What happens if a deployment fails during execution?
Marin-Deploy leverages Pulumi's state management and Kubernetes native health checks. If deployment fails, the previous healthy revision remains active. You can explicitly rollback to any prior revision using the --rollback flag followed by the revision number obtained from Deployment annotations, as implemented in lib/finelog/src/finelog/deploy/_k8s.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →