How to Debug Issues in Marin: A Complete Troubleshooting Guide

Use the Iris CLI to inspect scheduler state, query the controller SQLite database, profile running tasks, and roll back checkpoints when the Marin controller becomes unresponsive.

Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When jobs fail or the controller misbehaves, understanding how to debug issues in Marin requires navigating its distributed architecture and read-only introspection tools. This guide walks through the precise commands and source files you need to diagnose failures without risking state corruption.

Understanding the Debugging Architecture

Marin delegates execution across four distinct layers, each with its own debugging surface. Knowing which component owns your failure is the first step in any troubleshooting workflow.

  • Iris (job scheduler): Dispatches jobs, tracks task state in an on-VM SQLite database, and exposes RPC and CLI entry points defined in lib/iris/src/iris/__init__.py.
  • Levanter (JAX training): Handles data pipelines, sharding, and checkpointing for large models; reference lib/levanter/AGENTS.md for agent-specific diagnostics.
  • Zephyr (dataset processing): Manages Parquet reading, writers, and in-memory caching; see lib/zephyr/OPS.md for I/O debugging patterns.
  • Finelog (time-series metrics): Stores per-worker and per-task metrics separate from the controller DB, accessible via the namespaces defined in lib/finelog/README.md.

The Iris controller runs as a Docker container and persists its state to $STATE_DIR/db/controller.sqlite3, restoring from GCS checkpoints stored at $REMOTE/controller-state/. According to lib/iris/OPS.md, all Iris interactions are read-only by default, ensuring you can investigate safely before applying mutating commands.

Essential Debugging Commands

Verify Cluster Connectivity

Start every debugging session by confirming the controller is reachable and running the expected image version. Run:

iris --cluster=<NAME> cluster status

This command prints the current Git short-hash and health status. As noted in lib/iris/OPS.md, always verify the image hash matches your intended deployment before proceeding with destructive operations.

Inspect Scheduler State

To view queue length, resource constraints, and priority bands, query the scheduler directly:

iris rpc controller get-scheduler-state

This RPC returns the current allocation state and helps identify bottlenecked resources. For autoscaling decisions, use iris rpc controller get-autoscaler-status and look for the backoff_until_ms field to detect throttling.

Query the Controller Database

The controller stores job and task metadata in SQLite. Use the iris query command to inspect stuck jobs without SSH access:


# Count jobs by state

iris query "SELECT state, count(*) FROM jobs GROUP BY state"

# Export stuck tasks to CSV (state 5 = FAILED)

iris query -f csv "SELECT task_id, state FROM tasks WHERE state=5"

In lib/iris/OPS.md, state 5 corresponds to FAILED, while state 3 indicates RUNNING. Use these queries to identify jobs stuck in BUILDING or RUNNING states before pulling logs.

Analyze Job Logs

Once you identify a suspect job, extract specific error patterns without streaming the entire history:

iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"

Replace "Saving checkpoint" with "OOM" or "Error" to isolate Out-of-Memory kills or stack traces. The --no-tail flag fetches historical logs rather than following new output.

Profile Running Tasks

When performance degrades, generate CPU or memory flamegraphs without manual SSH or perf installation:


# Generate speedscope JSON for CPU analysis

iris process profile cpu -t /user/job/0

# Generate HTML flamegraph for memory

iris process profile mem -t /user/job/0

These commands, documented in lib/iris/OPS.md, write .speedscope.json or HTML files to your local machine, enabling offline analysis even when the cluster is behind IAP.

Advanced Recovery Techniques

Controller Checkpoint Rollback

If a deployment introduces instability, roll back both the container image and the database state:

iris cluster controller restart --rollback

This restores the previous image and its pre-deployment checkpoint from GCS. According to the source in lib/iris/OPS.md, checkpoints are stored at $REMOTE/controller-state/ and restored to $STATE_DIR/db/controller.sqlite3.

Handling Wedged Controllers

When the controller becomes unresponsive due to database corruption or bloat, never delete the SQLite file. Instead, follow the Checkpoint rollback procedure in lib/iris/OPS.md:

  1. Move the bloated database aside: mv $STATE_DIR/db/controller.sqlite3 $STATE_DIR/db/controller.sqlite3.bak
  2. Run download_checkpoint_to_local inside a temporary container to fetch the last known-good checkpoint from GCS
  3. Restart the controller to load the clean checkpoint

This workflow preserves historical metadata while recovering from disk-full or transaction-lock scenarios.

Automating Debug Workflows with Python

Automate repetitive investigations using the Iris CLI from Python scripts. Save this helper as debug_iris.py:

#!/usr/bin/env python3
import subprocess
import sys

def run(*cmd: str) -> str:
    """Run a shell command and return stdout."""
    result = subprocess.run(cmd, capture_output=True, text=True, check=False)
    if result.returncode != 0:
        print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
        sys.exit(1)
    return result.stdout.strip()

cluster = "marin"
job = "/user/example-job"

# Verify controller connectivity

print("🔎 Controller status")
print(run("iris", f"--cluster={cluster}", "cluster", "status"))

# Inspect scheduler state

print("\n📊 Scheduler state")
print(run("iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))

# Search logs for OOM patterns

print("\n🔎 Searching logs for OOM")
log = run(
    "iris", f"--cluster={cluster}", "job", "logs", job,
    "--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")

Make the script executable with chmod +x debug_iris.py and invoke it to standardize health checks across your team.

Summary

  • Always verify first: Run iris cluster status to confirm controller health and image hash before mutating operations.
  • Query before acting: Use iris query to inspect the SQLite database at $STATE_DIR/db/controller.sqlite3 to identify stuck jobs without restarting services.
  • Profile remotely: Generate .speedscope.json flamegraphs via iris process profile to diagnose performance issues without SSH access.
  • Rollback safely: Use iris cluster controller restart --rollback to restore both the container image and database checkpoint when deployments fail.
  • Never delete the DB: When the controller is wedged, move the SQLite file aside and restore from GCS checkpoints rather than deleting state.

Frequently Asked Questions

How do I check if the Marin controller is healthy?

Run iris --cluster=<NAME> cluster status to verify connectivity and view the current Git short-hash. A healthy controller returns promptly with image metadata; timeouts or connection errors indicate network issues or a crashed container.

What does state=5 mean in Marin job queries?

State 5 represents FAILED jobs in the controller SQLite schema. Use iris query "SELECT job_id, state FROM jobs WHERE state=5" to list failed jobs, then inspect logs with iris job logs /user/<job-name> to retrieve error details.

How can I profile a Marin task without SSH access?

Use the built-in RPC profiler: iris process profile cpu -t /user/job/0 generates a speedscope-compatible JSON file locally. This method works even when workers are behind IAP or in private Kubernetes clusters, unlike manual SSH-based profiling.

How do I safely restart the Marin controller?

First verify the current image hash with iris cluster status. Then run iris cluster controller restart to deploy the current checkout. Never restart without checking the hash, as this command ships the active Git state—deploying from a stale branch will roll back code changes. If the restart causes instability, immediately run iris cluster controller restart --rollback to restore the previous image and checkpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →