How to Debug Issues in Marin: A Complete Troubleshooting Guide
To debug issues in Marin, use the Iris CLI to inspect scheduler state and job logs, leverage built-in process profiling tools to capture CPU and memory flamegraphs, and restore from controller checkpoints when the database becomes corrupted.
Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When you need to debug issues in Marin, understanding the interaction between these layers is essential for efficient troubleshooting. This guide provides a systematic workflow based on the actual source code in the marin-community/marin repository, covering everything from read-only cluster inspection to safe controller recovery.
Understanding Marin's Debugging Architecture
Marin's debugging capabilities span four distinct layers, each with specific diagnostic entry points and state stores.
The Iris Controller Layer
The Iris controller runs as a Docker container on GCE or Kubernetes VMs and serves as the central nervous system for job orchestration. It maintains its state in $STATE_DIR/db/controller.sqlite3 and periodically backs up to GCS checkpoints at $REMOTE/controller-state/. According to the operations guide in lib/iris/OPS.md (lines 119-158), the controller exposes read-only RPCs such as iris rpc controller get-scheduler-state and SQL query interfaces via iris query for safe inspection without mutation risks.
Worker and Task State Management
Workers register through either the Iris worker daemon on GCE or the Kubernetes driver on CoreWeave. Autoscaling decisions and task transitions are recorded in the iris.worker and iris.task namespaces within the Finelog time-series metrics system (referenced in lib/iris/OPS.md lines 44-71). This separation allows you to debug resource allocation issues independently of the controller's SQLite state.
Step-by-Step Debugging Workflow
Follow this sequential checklist to isolate issues without risking cluster state. All Iris interactions are read-only by default, ensuring you can investigate safely before applying fixes.
1. Verify Cluster Connectivity
Start by confirming the controller is reachable and note the current deployment hash.
iris --cluster=<NAME> cluster status
This command shows the controller health and current image hash, which is critical before any restart operations (see Cluster Lifecycle in lib/iris/OPS.md lines 47-55).
2. Inspect Scheduler State
Check for queue bottlenecks and resource constraints.
iris rpc controller get-scheduler-state
Look for queue length, resource constraints, and priority band distributions in the output.
3. Query Job and Task State
Use SQL to identify stuck jobs (states like BUILDING or RUNNING that have hung):
# Count jobs by state
iris query "SELECT state, count(*) FROM jobs GROUP BY state"
# Find specific failed tasks
iris query -f csv "SELECT task_id, state FROM tasks WHERE state=3"
State 3 typically indicates FAILED status. Refer to the SQL Queries section of lib/iris/OPS.md (lines 119-140) for the full schema.
4. Analyze Job Logs
Pull complete logs for pattern matching without tailing:
iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"
This is particularly effective for finding OOM errors or checkpoint failures (see Log filtering in lib/iris/OPS.md lines 19-27).
5. Profile Running Processes
Capture CPU or memory profiles for any running task:
# CPU profiling
iris process profile cpu -t /user/job/0
# Memory profiling
iris process profile mem -t /user/job/0
These commands generate .speedscope.json files or HTML flamegraphs that you can analyze in Speedscope. See Process Inspection & Profiling in lib/iris/OPS.md (lines 71-84).
6. Create and Analyze Checkpoints
For offline analysis of controller state without impacting the live cluster:
# Create a GCS checkpoint
iris cluster controller checkpoint
# Download and query locally
sqlite3 /tmp/controller.sqlite3 "SELECT * FROM jobs WHERE state=5"
Never query the live controller database directly. Always work on a checkpoint copy (see Offline checkpoint analysis in lib/iris/OPS.md lines 28-36).
7. Restart the Controller
When you need to deploy fixed code:
iris cluster controller restart
Verify the new hash with iris cluster status to confirm the deployment succeeded (see Controller Restart in lib/iris/OPS.md lines 55-71).
8. Rollback Failed Deployments
If a restart introduces regressions, restore the previous state:
iris cluster controller restart --rollback
This restores both the previous container image and its pre-deployment checkpoint (see Rollback a controller deploy in lib/iris/OPS.md lines 103-110).
9. Recover Corrupted Controller State
When the controller becomes unresponsive due to database bloat:
- Move the bloated database aside (do not delete)
- Run
download_checkpoint_to_localinside a temporary container - Restore from the latest known-good checkpoint
Detailed steps are in the Checkpoint rollback section of lib/iris/OPS.md (lines 121-158).
10. Validate System Health
After changes, confirm controller stability:
uv run --package marin-iris --extra controller iris --cluster=<NAME> cluster status
uv run --package marin-iris --extra controller iris --cluster=<NAME> process profile cpu -t /system/controller
Automating Debugging with Python Scripts
Create a reusable script at debug_iris.py to automate common investigations:
#!/usr/bin/env python3
import subprocess
import sys
from pathlib import Path
def run(*cmd: str) -> str:
"""Run a shell command and return stdout."""
result = subprocess.run(cmd, capture_output=True, text=True, check=False)
if result.returncode != 0:
print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
return result.stdout.strip()
# 1️⃣ Verify controller connectivity
cluster = "marin"
print("🔎 Controller status")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
"iris", f"--cluster={cluster}", "cluster", "status"))
# 2️⃣ Pull recent scheduler snapshot
print("\n📊 Scheduler state")
print(run("uv", "run", "--package", "marin-iris", "--extra", "controller",
"iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))
# 3️⃣ Search a job's logs for OOM patterns
job = "/user/example-job"
print("\n🔎 Searching logs for OOM")
log = run(
"uv", "run", "--package", "marin-iris", "--extra", "controller",
"iris", f"--cluster={cluster}", "job", "logs", job,
"--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")
Make the script executable (chmod +x debug_iris.py) and run it to quickly verify cluster health.
Key Source Files for Debugging
Understanding these files in the marin-community/marin repository helps you trace issues to their root:
lib/iris/OPS.md— The primary operations guide containing the complete debugging checklist and RPC specifications.lib/iris/src/iris/__init__.py— Registers CLI entry points and controller initialization logic.lib/iris/src/iris/managed_thread.py— Implements the lightweight thread pool used by the controller; check here for threading issues.lib/iris/src/iris/env_resources.py— Parses resource flags (--cpu,--memory,--gpu) and validates allocations.lib/iris/src/iris/time_proto.py— Protocol buffer definitions for RPC timestamps.lib/zephyr/OPS.md— Diagnostic patterns for dataset readers, writers, and shuffling operations.lib/finelog/README.md— Documentation for the metrics storage system used by workers and tasks.scripts/iris/rollout_controllers.py— Automation script for safe controller rollouts.
Critical Debugging Scenarios and Solutions
Controller Database Corruption
When the SQLite database at $STATE_DIR/db/controller.sqlite3 grows too large or becomes locked, the controller may hang. Do not delete the database. Instead, follow the checkpoint rollback procedure in lib/iris/OPS.md (lines 121-158) to move the bloated file aside and restore from GCS.
Image Architecture Mismatches
If GPU pods die immediately upon startup, verify the container architecture matches the node hardware. Query the container manifest using curl as shown in the Image architecture mismatch section of lib/iris/OPS.md (lines 84-91) to confirm ARM64 vs. AMD64 compatibility.
Stuck Jobs and Tasks
Jobs stuck in BUILDING or RUNNING states often indicate resource starvation or dependency failures. Use iris query to identify stuck tasks by state ID, then pull logs with iris job logs to find the specific error pattern. For tasks that consume excessive resources, use iris process profile to capture flamegraphs before preemption.
Summary
- Use read-only commands first: Start with
iris cluster status,iris query, andiris job logsto investigate without risk. - Profile before restarting: Capture CPU and memory profiles using
iris process profileto preserve debugging data from failing tasks. - Checkpoint safely: Create GCS checkpoints before destructive operations, and always work on copies of the SQLite database, never the live controller state.
- Rollback carefully: Use
iris cluster controller restart --rollbackto revert both code and state when deployments fail. - Consult
lib/iris/OPS.md: This file contains the authoritative source for all debugging procedures referenced in this guide.
Frequently Asked Questions
How do I debug a Marin job that is stuck in the RUNNING state?
Use iris query "SELECT task_id, state FROM tasks WHERE state=2" to identify the specific task ID (state 2 typically indicates RUNNING), then retrieve logs with iris job logs /user/<job-name> --max-lines 100000 --substring "error" to find failure patterns. If the task is consuming excessive resources, run iris process profile cpu -t /user/job/<task-id> to generate a flamegraph and identify bottlenecks.
Where does Marin store controller state for debugging purposes?
The Iris controller stores its state in a SQLite database at $STATE_DIR/db/controller.sqlite3 on the controller VM, with periodic backups written to GCS at $REMOTE/controller-state/. You can download these checkpoints using iris cluster controller checkpoint and analyze them locally with standard SQLite tools without impacting the live cluster.
What is the safest way to restart the Marin controller if it becomes unresponsive?
First, create a checkpoint using iris cluster controller checkpoint. Then execute iris cluster controller restart to deploy the current Git checkout. After restart, immediately verify the new image hash with iris cluster status. If the restart worsens the issue, use iris cluster controller restart --rollback to restore the previous image and its pre-deployment checkpoint state.
How can I automate debugging tasks in Marin without manual CLI typing?
Create a Python script that uses subprocess.run() to execute uv run --package marin-iris --extra controller iris commands, as shown in the automation example above. This allows you to script connectivity checks, log searches for specific error patterns like "OOM", and scheduler state snapshots. Always use --dry-run flags when available in automation to preview mutations before applying them.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →