# How to Debug Issues in Marin: A Complete Troubleshooting Guide

> Debug Marin issues effectively. Use Iris CLI to inspect scheduler, query the controller database, profile tasks, and roll back checkpoints when Marin becomes unresponsive.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-27

---

**Use the Iris CLI to inspect scheduler state, query the controller SQLite database, profile running tasks, and roll back checkpoints when the Marin controller becomes unresponsive.**

Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When jobs fail or the controller misbehaves, understanding how to debug issues in Marin requires navigating its distributed architecture and read-only introspection tools. This guide walks through the precise commands and source files you need to diagnose failures without risking state corruption.

## Understanding the Debugging Architecture

Marin delegates execution across four distinct layers, each with its own debugging surface. Knowing which component owns your failure is the first step in any troubleshooting workflow.

- **Iris** (job scheduler): Dispatches jobs, tracks task state in an on-VM SQLite database, and exposes RPC and CLI entry points defined in [`lib/iris/src/iris/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/__init__.py).
- **Levanter** (JAX training): Handles data pipelines, sharding, and checkpointing for large models; reference [`lib/levanter/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/levanter/AGENTS.md) for agent-specific diagnostics.
- **Zephyr** (dataset processing): Manages Parquet reading, writers, and in-memory caching; see [`lib/zephyr/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/zephyr/OPS.md) for I/O debugging patterns.
- **Finelog** (time-series metrics): Stores per-worker and per-task metrics separate from the controller DB, accessible via the namespaces defined in [`lib/finelog/README.md`](https://github.com/marin-community/marin/blob/main/lib/finelog/README.md).

The Iris controller runs as a Docker container and persists its state to `$STATE_DIR/db/controller.sqlite3`, restoring from GCS checkpoints stored at `$REMOTE/controller-state/`. According to [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), all Iris interactions are **read-only by default**, ensuring you can investigate safely before applying mutating commands.

## Essential Debugging Commands

### Verify Cluster Connectivity

Start every debugging session by confirming the controller is reachable and running the expected image version. Run:

```bash
iris --cluster=<NAME> cluster status

```

This command prints the current Git short-hash and health status. As noted in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), always verify the image hash matches your intended deployment before proceeding with destructive operations.

### Inspect Scheduler State

To view queue length, resource constraints, and priority bands, query the scheduler directly:

```bash
iris rpc controller get-scheduler-state

```

This RPC returns the current allocation state and helps identify bottlenecked resources. For autoscaling decisions, use `iris rpc controller get-autoscaler-status` and look for the `backoff_until_ms` field to detect throttling.

### Query the Controller Database

The controller stores job and task metadata in SQLite. Use the `iris query` command to inspect stuck jobs without SSH access:

```bash

# Count jobs by state

iris query "SELECT state, count(*) FROM jobs GROUP BY state"

# Export stuck tasks to CSV (state 5 = FAILED)

iris query -f csv "SELECT task_id, state FROM tasks WHERE state=5"

```

In [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), state 5 corresponds to **FAILED**, while state 3 indicates **RUNNING**. Use these queries to identify jobs stuck in `BUILDING` or `RUNNING` states before pulling logs.

### Analyze Job Logs

Once you identify a suspect job, extract specific error patterns without streaming the entire history:

```bash
iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"

```

Replace `"Saving checkpoint"` with `"OOM"` or `"Error"` to isolate Out-of-Memory kills or stack traces. The `--no-tail` flag fetches historical logs rather than following new output.

### Profile Running Tasks

When performance degrades, generate CPU or memory flamegraphs without manual SSH or `perf` installation:

```bash

# Generate speedscope JSON for CPU analysis

iris process profile cpu -t /user/job/0

# Generate HTML flamegraph for memory

iris process profile mem -t /user/job/0

```

These commands, documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), write [`.speedscope.json`](https://github.com/marin-community/marin/blob/main/.speedscope.json) or HTML files to your local machine, enabling offline analysis even when the cluster is behind IAP.

## Advanced Recovery Techniques

### Controller Checkpoint Rollback

If a deployment introduces instability, roll back both the container image and the database state:

```bash
iris cluster controller restart --rollback

```

This restores the previous image and its pre-deployment checkpoint from GCS. According to the source in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), checkpoints are stored at `$REMOTE/controller-state/` and restored to `$STATE_DIR/db/controller.sqlite3`.

### Handling Wedged Controllers

When the controller becomes unresponsive due to database corruption or bloat, never delete the SQLite file. Instead, follow the **Checkpoint rollback** procedure in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md):

1. Move the bloated database aside: `mv $STATE_DIR/db/controller.sqlite3 $STATE_DIR/db/controller.sqlite3.bak`
2. Run `download_checkpoint_to_local` inside a temporary container to fetch the last known-good checkpoint from GCS
3. Restart the controller to load the clean checkpoint

This workflow preserves historical metadata while recovering from disk-full or transaction-lock scenarios.

## Automating Debug Workflows with Python

Automate repetitive investigations using the Iris CLI from Python scripts. Save this helper as [`debug_iris.py`](https://github.com/marin-community/marin/blob/main/debug_iris.py):

```python
#!/usr/bin/env python3
import subprocess
import sys

def run(*cmd: str) -> str:
    """Run a shell command and return stdout."""
    result = subprocess.run(cmd, capture_output=True, text=True, check=False)
    if result.returncode != 0:
        print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
        sys.exit(1)
    return result.stdout.strip()

cluster = "marin"
job = "/user/example-job"

# Verify controller connectivity

print("🔎 Controller status")
print(run("iris", f"--cluster={cluster}", "cluster", "status"))

# Inspect scheduler state

print("\n📊 Scheduler state")
print(run("iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))

# Search logs for OOM patterns

print("\n🔎 Searching logs for OOM")
log = run(
    "iris", f"--cluster={cluster}", "job", "logs", job,
    "--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")

```

Make the script executable with `chmod +x debug_iris.py` and invoke it to standardize health checks across your team.

## Summary

- **Always verify first**: Run `iris cluster status` to confirm controller health and image hash before mutating operations.
- **Query before acting**: Use `iris query` to inspect the SQLite database at `$STATE_DIR/db/controller.sqlite3` to identify stuck jobs without restarting services.
- **Profile remotely**: Generate [`.speedscope.json`](https://github.com/marin-community/marin/blob/main/.speedscope.json) flamegraphs via `iris process profile` to diagnose performance issues without SSH access.
- **Rollback safely**: Use `iris cluster controller restart --rollback` to restore both the container image and database checkpoint when deployments fail.
- **Never delete the DB**: When the controller is wedged, move the SQLite file aside and restore from GCS checkpoints rather than deleting state.

## Frequently Asked Questions

### How do I check if the Marin controller is healthy?

Run `iris --cluster=<NAME> cluster status` to verify connectivity and view the current Git short-hash. A healthy controller returns promptly with image metadata; timeouts or connection errors indicate network issues or a crashed container.

### What does state=5 mean in Marin job queries?

State 5 represents **FAILED** jobs in the controller SQLite schema. Use `iris query "SELECT job_id, state FROM jobs WHERE state=5"` to list failed jobs, then inspect logs with `iris job logs /user/<job-name>` to retrieve error details.

### How can I profile a Marin task without SSH access?

Use the built-in RPC profiler: `iris process profile cpu -t /user/job/0` generates a speedscope-compatible JSON file locally. This method works even when workers are behind IAP or in private Kubernetes clusters, unlike manual SSH-based profiling.

### How do I safely restart the Marin controller?

First verify the current image hash with `iris cluster status`. Then run `iris cluster controller restart` to deploy the current checkout. **Never restart without checking the hash**, as this command ships the active Git state—deploying from a stale branch will roll back code changes. If the restart causes instability, immediately run `iris cluster controller restart --rollback` to restore the previous image and checkpoint.