# How to Debug Issues in Marin: A Complete Troubleshooting Guide

> Troubleshoot Marin issues effectively. Learn to debug with Iris CLI, analyze flamegraphs, and restore from controller checkpoints. Master Marin troubleshooting today.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**To debug issues in Marin, use the Iris CLI to inspect scheduler state and job logs, leverage built-in process profiling tools to capture CPU and memory flamegraphs, and restore from controller checkpoints when the database becomes corrupted.**

Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When you need to debug issues in Marin, understanding the interaction between these layers is essential for efficient troubleshooting. This guide provides a systematic workflow based on the actual source code in the `marin-community/marin` repository, covering everything from read-only cluster inspection to safe controller recovery.

## Understanding Marin's Debugging Architecture

Marin's debugging capabilities span four distinct layers, each with specific diagnostic entry points and state stores.

### The Iris Controller Layer

The Iris controller runs as a Docker container on GCE or Kubernetes VMs and serves as the central nervous system for job orchestration. It maintains its state in `$STATE_DIR/db/controller.sqlite3` and periodically backs up to GCS checkpoints at `$REMOTE/controller-state/`. According to the operations guide in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 119-158), the controller exposes read-only RPCs such as `iris rpc controller get-scheduler-state` and SQL query interfaces via `iris query` for safe inspection without mutation risks.

### Worker and Task State Management

Workers register through either the Iris worker daemon on GCE or the Kubernetes driver on CoreWeave. Autoscaling decisions and task transitions are recorded in the `iris.worker` and `iris.task` namespaces within the Finelog time-series metrics system (referenced in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 44-71). This separation allows you to debug resource allocation issues independently of the controller's SQLite state.

## Step-by-Step Debugging Workflow

Follow this sequential checklist to isolate issues without risking cluster state. All Iris interactions are **read-only by default**, ensuring you can investigate safely before applying fixes.

### 1. Verify Cluster Connectivity

Start by confirming the controller is reachable and note the current deployment hash.

```bash
iris --cluster=<NAME> cluster status

```

This command shows the controller health and current image hash, which is critical before any restart operations (see *Cluster Lifecycle* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 47-55).

### 2. Inspect Scheduler State

Check for queue bottlenecks and resource constraints.

```bash
iris rpc controller get-scheduler-state

```

Look for queue length, resource constraints, and priority band distributions in the output.

### 3. Query Job and Task State

Use SQL to identify stuck jobs (states like `BUILDING` or `RUNNING` that have hung):

```bash

# Count jobs by state

iris query "SELECT state, count(*) FROM jobs GROUP BY state"

# Find specific failed tasks

iris query -f csv "SELECT task_id, state FROM tasks WHERE state=3"

```

State `3` typically indicates `FAILED` status. Refer to the *SQL Queries* section of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 119-140) for the full schema.

### 4. Analyze Job Logs

Pull complete logs for pattern matching without tailing:

```bash
iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"

```

This is particularly effective for finding OOM errors or checkpoint failures (see *Log filtering* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 19-27).

### 5. Profile Running Processes

Capture CPU or memory profiles for any running task:

```bash

# CPU profiling

iris process profile cpu -t /user/job/0

# Memory profiling

iris process profile mem -t /user/job/0

```

These commands generate [`.speedscope.json`](https://github.com/marin-community/marin/blob/main/.speedscope.json) files or HTML flamegraphs that you can analyze in Speedscope. See *Process Inspection & Profiling* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 71-84).

### 6. Create and Analyze Checkpoints

For offline analysis of controller state without impacting the live cluster:

```bash

# Create a GCS checkpoint

iris cluster controller checkpoint

# Download and query locally

sqlite3 /tmp/controller.sqlite3 "SELECT * FROM jobs WHERE state=5"

```

**Never query the live controller database directly.** Always work on a checkpoint copy (see *Offline checkpoint analysis* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 28-36).

### 7. Restart the Controller

When you need to deploy fixed code:

```bash
iris cluster controller restart

```

 Verify the new hash with `iris cluster status` to confirm the deployment succeeded (see *Controller Restart* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 55-71).

### 8. Rollback Failed Deployments

If a restart introduces regressions, restore the previous state:

```bash
iris cluster controller restart --rollback

```

This restores both the previous container image and its pre-deployment checkpoint (see *Rollback a controller deploy* in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 103-110).

### 9. Recover Corrupted Controller State

When the controller becomes unresponsive due to database bloat:

1. Move the bloated database aside (do not delete)
2. Run `download_checkpoint_to_local` inside a temporary container
3. Restore from the latest known-good checkpoint

Detailed steps are in the *Checkpoint rollback* section of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 121-158).

### 10. Validate System Health

After changes, confirm controller stability:

```bash
uv run --package marin-iris --extra controller iris --cluster=<NAME> cluster status
uv run --package marin-iris --extra controller iris --cluster=<NAME> process profile cpu -t /system/controller

```

## Automating Debugging with Python Scripts

Create a reusable script at [`debug_iris.py`](https://github.com/marin-community/marin/blob/main/debug_iris.py) to automate common investigations:

```python
#!/usr/bin/env python3
import subprocess
import sys
from pathlib import Path

def run(*cmd: str) -> str:
    """Run a shell command and return stdout."""
    result = subprocess.run(cmd, capture_output=True, text=True, check=False)
    if result.returncode != 0:
        print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
    return result.stdout.strip()

# 1️⃣ Verify controller connectivity

cluster = "marin"
print("🔎 Controller status")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
          "iris", f"--cluster={cluster}", "cluster", "status"))

# 2️⃣ Pull recent scheduler snapshot

print("\n📊 Scheduler state")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
          "iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))

# 3️⃣ Search a job's logs for OOM patterns

job = "/user/example-job"
print("\n🔎 Searching logs for OOM")
log = run(
    "uv", "run", "--package", "marin-iris", "--extra", "controller",
    "iris", f"--cluster={cluster}", "job", "logs", job,
    "--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")

```

Make the script executable (`chmod +x debug_iris.py`) and run it to quickly verify cluster health.

## Key Source Files for Debugging

Understanding these files in the `marin-community/marin` repository helps you trace issues to their root:

- **[`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md)** — The primary operations guide containing the complete debugging checklist and RPC specifications.
- **[`lib/iris/src/iris/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/__init__.py)** — Registers CLI entry points and controller initialization logic.
- **[`lib/iris/src/iris/managed_thread.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/managed_thread.py)** — Implements the lightweight thread pool used by the controller; check here for threading issues.
- **[`lib/iris/src/iris/env_resources.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/env_resources.py)** — Parses resource flags (`--cpu`, `--memory`, `--gpu`) and validates allocations.
- **[`lib/iris/src/iris/time_proto.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/time_proto.py)** — Protocol buffer definitions for RPC timestamps.
- **[`lib/zephyr/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/zephyr/OPS.md)** — Diagnostic patterns for dataset readers, writers, and shuffling operations.
- **[`lib/finelog/README.md`](https://github.com/marin-community/marin/blob/main/lib/finelog/README.md)** — Documentation for the metrics storage system used by workers and tasks.
- **[`scripts/iris/rollout_controllers.py`](https://github.com/marin-community/marin/blob/main/scripts/iris/rollout_controllers.py)** — Automation script for safe controller rollouts.

## Critical Debugging Scenarios and Solutions

### Controller Database Corruption

When the SQLite database at `$STATE_DIR/db/controller.sqlite3` grows too large or becomes locked, the controller may hang. **Do not delete the database.** Instead, follow the checkpoint rollback procedure in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 121-158) to move the bloated file aside and restore from GCS.

### Image Architecture Mismatches

If GPU pods die immediately upon startup, verify the container architecture matches the node hardware. Query the container manifest using `curl` as shown in the *Image architecture mismatch* section of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) (lines 84-91) to confirm ARM64 vs. AMD64 compatibility.

### Stuck Jobs and Tasks

Jobs stuck in `BUILDING` or `RUNNING` states often indicate resource starvation or dependency failures. Use `iris query` to identify stuck tasks by state ID, then pull logs with `iris job logs` to find the specific error pattern. For tasks that consume excessive resources, use `iris process profile` to capture flamegraphs before preemption.

## Summary

- **Use read-only commands first:** Start with `iris cluster status`, `iris query`, and `iris job logs` to investigate without risk.
- **Profile before restarting:** Capture CPU and memory profiles using `iris process profile` to preserve debugging data from failing tasks.
- **Checkpoint safely:** Create GCS checkpoints before destructive operations, and always work on copies of the SQLite database, never the live controller state.
- **Rollback carefully:** Use `iris cluster controller restart --rollback` to revert both code and state when deployments fail.
- **Consult [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md):** This file contains the authoritative source for all debugging procedures referenced in this guide.

## Frequently Asked Questions

### How do I debug a Marin job that is stuck in the RUNNING state?

Use `iris query "SELECT task_id, state FROM tasks WHERE state=2"` to identify the specific task ID (state 2 typically indicates RUNNING), then retrieve logs with `iris job logs /user/<job-name> --max-lines 100000 --substring "error"` to find failure patterns. If the task is consuming excessive resources, run `iris process profile cpu -t /user/job/<task-id>` to generate a flamegraph and identify bottlenecks.

### Where does Marin store controller state for debugging purposes?

The Iris controller stores its state in a SQLite database at `$STATE_DIR/db/controller.sqlite3` on the controller VM, with periodic backups written to GCS at `$REMOTE/controller-state/`. You can download these checkpoints using `iris cluster controller checkpoint` and analyze them locally with standard SQLite tools without impacting the live cluster.

### What is the safest way to restart the Marin controller if it becomes unresponsive?

First, create a checkpoint using `iris cluster controller checkpoint`. Then execute `iris cluster controller restart` to deploy the current Git checkout. After restart, immediately verify the new image hash with `iris cluster status`. If the restart worsens the issue, use `iris cluster controller restart --rollback` to restore the previous image and its pre-deployment checkpoint state.

### How can I automate debugging tasks in Marin without manual CLI typing?

Create a Python script that uses `subprocess.run()` to execute `uv run --package marin-iris --extra controller iris` commands, as shown in the automation example above. This allows you to script connectivity checks, log searches for specific error patterns like "OOM", and scheduler state snapshots. Always use `--dry-run` flags when available in automation to preview mutations before applying them.