How to Deploy the Iris Controller to CoreWeave Kubernetes: A Complete Guide
Deploy the Iris controller to CoreWeave Kubernetes by configuring a coreweave platform section in your cluster YAML, building the controller image with the rollout_controllers.py script, and applying the generated Kubernetes manifests to the target namespace.
The Iris controller is the central component in the marin-community/marin repository that runs inside a Kubernetes pod and communicates with the Iris service via RPC. When targeting a CoreWeave cluster, the deployment process follows standard Kubernetes patterns but requires CoreWeave-specific configuration values to locate the cluster region and object storage endpoints. This guide walks through the exact configuration schema, build commands, and verification steps needed to run the controller on CoreWeave infrastructure.
Configure the CoreWeave Platform Settings
Before deploying, you must define a coreweave platform configuration in your Iris cluster YAML file. This configuration tells the controller how to authenticate with CoreWeave and where to store artifacts.
Define the Platform Section
Create or edit an Iris cluster configuration file (for example, config/coreweave-production.yaml) and add the coreweave block under the platform key. The schema for this block is defined in lib/iris/src/iris/cluster/config.py at lines 19-27, which implements the CoreweavePlatformConfig dataclass.
platform:
coreweave:
region: "US-EAST-02A" # Required: CoreWeave region identifier
namespace: "iris" # Required: Target Kubernetes namespace
kubeconfig_path: "~/.kube/coreweave-iris" # Required: Path to CoreWeave kubeconfig
kube_context: "" # Optional: Explicit kubectl context name
object_storage_endpoint: "https://object.example.coreweave.com"
external_object_storage_endpoint: "" # Optional: Falls back to object_storage_endpoint if empty
The controller uses these values to determine which CoreWeave region to target and which namespace should host the controller pod.
Configure S3-Compatible Storage
The Iris controller requires a configured S3-compatible client before making any RPC calls that touch object storage. According to the source code in lib/rigging/src/rigging/filesystem/s3_compat.py at lines 176-182, you must call configure_coreweave_s3() to initialize the process-wide storage client.
When using the automated rollout script, this happens automatically. For manual debugging, invoke it directly:
from rigging.filesystem.s3_compat import configure_coreweave_s3
configure_coreweave_s3()
Build and Deploy the Controller
The Marin repository includes automation scripts that handle Docker builds, image pushes, and Kubernetes manifest generation for CoreWeave deployments.
Build and Push the Docker Image
Use the coreweave-controller command in scripts/iris/rollout_controllers.py to build the controller image. This command uses the Dockerfile located at docker/iris-controller and validates your configuration against the require_coreweave_platform helper defined in scripts/iris/dev_gpu.py (lines 131-153).
Run the following command to build and push the image to the CoreWeave registry:
uv run scripts/iris/rollout_controllers.py coreweave-controller \
--config config/coreweave-production.yaml \
--push
The --push flag ensures the image is uploaded to the registry configured for your specified CoreWeave region.
Apply Kubernetes Manifests
After the image build completes, the rollout script automatically generates a temporary Kubernetes manifest containing a Deployment and a Service. It applies these manifests using the kubeconfig path specified in your configuration. By default, the controller listens on port 10000 within the namespace you defined (typically iris).
You can verify the applied resources with standard kubectl commands:
kubectl --kubeconfig ~/.kube/coreweave-iris get pods -n iris
Verify and Debug the Deployment
Once the manifests are applied, the rollout script performs automated health checks to confirm the controller is serving RPC requests.
Health Check Verification
The script waits for the controller pod to reach the Ready state, then executes an RPC health check (iris client version) to confirm functionality. If the check fails, the script aborts and outputs the pod logs for debugging.
To run verification independently:
uv run scripts/iris/rollout_controllers.py coreweave-controller \
--config config/coreweave-production.yaml \
--verify
Local Port Forwarding for Debugging
For local development or debugging, you can establish a port-forward tunnel to the controller pod without exposing it externally. The script implements this via _start_coreweave_port_forward at lines 1050-1070 of rollout_controllers.py.
Execute the following to forward the controller port to your local machine:
uv run scripts/iris/rollout_controllers.py coreweave-controller \
--config config/coreweave-production.yaml \
--port-forward
This runs kubectl port-forward against the identified controller pod, allowing local RPC clients to communicate with the service.
Complete Deployment Example
The following example demonstrates the full workflow from configuration to verification:
# 1. Create the CoreWeave platform configuration
cat > config/coreweave-production.yaml <<'EOF'
platform:
coreweave:
region: "US-EAST-02A"
namespace: "iris"
kubeconfig_path: "~/.kube/coreweave-iris"
object_storage_endpoint: "https://object.example.coreweave.com"
EOF
# 2. Build, push, and deploy the controller
uv run scripts/iris/rollout_controllers.py coreweave-controller \
--config config/coreweave-production.yaml \
--push
# 3. Verify the deployment is healthy
uv run scripts/iris/rollout_controllers.py coreweave-controller \
--config config/coreweave-production.yaml \
--verify
The example configuration above matches the schema used in the CI smoke tests located at config/ci-coreweave-gpu-smoke.yaml.
Summary
- Configuration Schema: Define the
coreweaveplatform section in your cluster YAML following theCoreweavePlatformConfigstructure inlib/iris/src/iris/cluster/config.py. - Build Automation: Use
scripts/iris/rollout_controllers.pywith thecoreweave-controllersubcommand to handle Docker builds and registry pushes. - Storage Setup: The controller automatically calls
configure_coreweave_s3()fromlib/rigging/src/rigging/filesystem/s3_compat.pyto initialize S3-compatible storage access. - Deployment: The script generates Kubernetes
DeploymentandServicemanifests, applying them to the namespace specified in your config. - Debugging: Use the
--port-forwardflag to create local tunnels for RPC testing, or check logs when health checks fail.
Frequently Asked Questions
What is the Iris controller in the Marin project?
The Iris controller is the central orchestration component that runs as a Kubernetes pod within the Marin machine learning infrastructure. It handles RPC requests from the Iris service and manages job execution, object storage interactions, and cluster resource coordination on CoreWeave GPUs.
How do I configure the CoreWeave object storage endpoint?
Set the object_storage_endpoint field in the platform.coreweave section of your cluster configuration YAML. This value should point to the S3-compatible endpoint provided by CoreWeave for your region. The controller uses this endpoint via the configure_coreweave_s3() helper to store and retrieve training artifacts and model checkpoints.
Can I deploy the controller to a different Kubernetes context?
Yes. While the kubeconfig_path field specifies the configuration file location, you can optionally set the kube_context field to target a specific context within that file. If left empty, the script uses the current default context from the specified kubeconfig.
How do I troubleshoot a failed Iris controller deployment?
Run the rollout script with the verification flag to trigger the health check loop, which prints pod logs automatically if the controller fails to become ready. You can also manually inspect the pod status using kubectl logs with the kubeconfig path specified in your Iris configuration, or use the --port-forward option to test RPC connectivity directly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →