How to Debug Issues in Marin: A Complete Troubleshooting Guide

To debug issues in Marin, use the Iris CLI to inspect scheduler state and job logs, leverage built-in process profiling tools to capture CPU and memory flamegraphs, and restore from controller checkpoints when the database becomes corrupted.

Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When you need to debug issues in Marin, understanding the interaction between these layers is essential for efficient troubleshooting. This guide provides a systematic workflow based on the actual source code in the marin-community/marin repository, covering everything from read-only cluster inspection to safe controller recovery.

Understanding Marin's Debugging Architecture

Marin's debugging capabilities span four distinct layers, each with specific diagnostic entry points and state stores.

The Iris Controller Layer

The Iris controller runs as a Docker container on GCE or Kubernetes VMs and serves as the central nervous system for job orchestration. It maintains its state in $STATE_DIR/db/controller.sqlite3 and periodically backs up to GCS checkpoints at $REMOTE/controller-state/. According to the operations guide in lib/iris/OPS.md (lines 119-158), the controller exposes read-only RPCs such as iris rpc controller get-scheduler-state and SQL query interfaces via iris query for safe inspection without mutation risks.

Worker and Task State Management

Workers register through either the Iris worker daemon on GCE or the Kubernetes driver on CoreWeave. Autoscaling decisions and task transitions are recorded in the iris.worker and iris.task namespaces within the Finelog time-series metrics system (referenced in lib/iris/OPS.md lines 44-71). This separation allows you to debug resource allocation issues independently of the controller's SQLite state.

Step-by-Step Debugging Workflow

Follow this sequential checklist to isolate issues without risking cluster state. All Iris interactions are read-only by default, ensuring you can investigate safely before applying fixes.

1. Verify Cluster Connectivity

Start by confirming the controller is reachable and note the current deployment hash.

iris --cluster=<NAME> cluster status

This command shows the controller health and current image hash, which is critical before any restart operations (see Cluster Lifecycle in lib/iris/OPS.md lines 47-55).

2. Inspect Scheduler State

Check for queue bottlenecks and resource constraints.

iris rpc controller get-scheduler-state

Look for queue length, resource constraints, and priority band distributions in the output.

3. Query Job and Task State

Use SQL to identify stuck jobs (states like BUILDING or RUNNING that have hung):


# Count jobs by state

iris query "SELECT state, count(*) FROM jobs GROUP BY state"

# Find specific failed tasks

iris query -f csv "SELECT task_id, state FROM tasks WHERE state=3"

State 3 typically indicates FAILED status. Refer to the SQL Queries section of lib/iris/OPS.md (lines 119-140) for the full schema.

4. Analyze Job Logs

Pull complete logs for pattern matching without tailing:

iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"

This is particularly effective for finding OOM errors or checkpoint failures (see Log filtering in lib/iris/OPS.md lines 19-27).

5. Profile Running Processes

Capture CPU or memory profiles for any running task:


# CPU profiling

iris process profile cpu -t /user/job/0

# Memory profiling

iris process profile mem -t /user/job/0

These commands generate .speedscope.json files or HTML flamegraphs that you can analyze in Speedscope. See Process Inspection & Profiling in lib/iris/OPS.md (lines 71-84).

6. Create and Analyze Checkpoints

For offline analysis of controller state without impacting the live cluster:


# Create a GCS checkpoint

iris cluster controller checkpoint

# Download and query locally

sqlite3 /tmp/controller.sqlite3 "SELECT * FROM jobs WHERE state=5"

Never query the live controller database directly. Always work on a checkpoint copy (see Offline checkpoint analysis in lib/iris/OPS.md lines 28-36).

7. Restart the Controller

When you need to deploy fixed code:

iris cluster controller restart

Verify the new hash with iris cluster status to confirm the deployment succeeded (see Controller Restart in lib/iris/OPS.md lines 55-71).

8. Rollback Failed Deployments

If a restart introduces regressions, restore the previous state:

iris cluster controller restart --rollback

This restores both the previous container image and its pre-deployment checkpoint (see Rollback a controller deploy in lib/iris/OPS.md lines 103-110).

9. Recover Corrupted Controller State

When the controller becomes unresponsive due to database bloat:

  1. Move the bloated database aside (do not delete)
  2. Run download_checkpoint_to_local inside a temporary container
  3. Restore from the latest known-good checkpoint

Detailed steps are in the Checkpoint rollback section of lib/iris/OPS.md (lines 121-158).

10. Validate System Health

After changes, confirm controller stability:

uv run --package marin-iris --extra controller iris --cluster=<NAME> cluster status
uv run --package marin-iris --extra controller iris --cluster=<NAME> process profile cpu -t /system/controller

Automating Debugging with Python Scripts

Create a reusable script at debug_iris.py to automate common investigations:

#!/usr/bin/env python3
import subprocess
import sys
from pathlib import Path

def run(*cmd: str) -> str:
    """Run a shell command and return stdout."""
    result = subprocess.run(cmd, capture_output=True, text=True, check=False)
    if result.returncode != 0:
        print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
    return result.stdout.strip()

# 1️⃣ Verify controller connectivity

cluster = "marin"
print("🔎 Controller status")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
          "iris", f"--cluster={cluster}", "cluster", "status"))

# 2️⃣ Pull recent scheduler snapshot

print("\n📊 Scheduler state")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
          "iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))

# 3️⃣ Search a job's logs for OOM patterns

job = "/user/example-job"
print("\n🔎 Searching logs for OOM")
log = run(
    "uv", "run", "--package", "marin-iris", "--extra", "controller",
    "iris", f"--cluster={cluster}", "job", "logs", job,
    "--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")

Make the script executable (chmod +x debug_iris.py) and run it to quickly verify cluster health.

Key Source Files for Debugging

Understanding these files in the marin-community/marin repository helps you trace issues to their root:

Critical Debugging Scenarios and Solutions

Controller Database Corruption

When the SQLite database at $STATE_DIR/db/controller.sqlite3 grows too large or becomes locked, the controller may hang. Do not delete the database. Instead, follow the checkpoint rollback procedure in lib/iris/OPS.md (lines 121-158) to move the bloated file aside and restore from GCS.

Image Architecture Mismatches

If GPU pods die immediately upon startup, verify the container architecture matches the node hardware. Query the container manifest using curl as shown in the Image architecture mismatch section of lib/iris/OPS.md (lines 84-91) to confirm ARM64 vs. AMD64 compatibility.

Stuck Jobs and Tasks

Jobs stuck in BUILDING or RUNNING states often indicate resource starvation or dependency failures. Use iris query to identify stuck tasks by state ID, then pull logs with iris job logs to find the specific error pattern. For tasks that consume excessive resources, use iris process profile to capture flamegraphs before preemption.

Summary

  • Use read-only commands first: Start with iris cluster status, iris query, and iris job logs to investigate without risk.
  • Profile before restarting: Capture CPU and memory profiles using iris process profile to preserve debugging data from failing tasks.
  • Checkpoint safely: Create GCS checkpoints before destructive operations, and always work on copies of the SQLite database, never the live controller state.
  • Rollback carefully: Use iris cluster controller restart --rollback to revert both code and state when deployments fail.
  • Consult lib/iris/OPS.md: This file contains the authoritative source for all debugging procedures referenced in this guide.

Frequently Asked Questions

How do I debug a Marin job that is stuck in the RUNNING state?

Use iris query "SELECT task_id, state FROM tasks WHERE state=2" to identify the specific task ID (state 2 typically indicates RUNNING), then retrieve logs with iris job logs /user/<job-name> --max-lines 100000 --substring "error" to find failure patterns. If the task is consuming excessive resources, run iris process profile cpu -t /user/job/<task-id> to generate a flamegraph and identify bottlenecks.

Where does Marin store controller state for debugging purposes?

The Iris controller stores its state in a SQLite database at $STATE_DIR/db/controller.sqlite3 on the controller VM, with periodic backups written to GCS at $REMOTE/controller-state/. You can download these checkpoints using iris cluster controller checkpoint and analyze them locally with standard SQLite tools without impacting the live cluster.

What is the safest way to restart the Marin controller if it becomes unresponsive?

First, create a checkpoint using iris cluster controller checkpoint. Then execute iris cluster controller restart to deploy the current Git checkout. After restart, immediately verify the new image hash with iris cluster status. If the restart worsens the issue, use iris cluster controller restart --rollback to restore the previous image and its pre-deployment checkpoint state.

How can I automate debugging tasks in Marin without manual CLI typing?

Create a Python script that uses subprocess.run() to execute uv run --package marin-iris --extra controller iris commands, as shown in the automation example above. This allows you to script connectivity checks, log searches for specific error patterns like "OOM", and scheduler state snapshots. Always use --dry-run flags when available in automation to preview mutations before applying them.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →