How to Debug Issues in Marin: A Complete Troubleshooting Guide
Use the Iris CLI to inspect scheduler state, query the controller SQLite database, profile running tasks, and roll back checkpoints when the Marin controller becomes unresponsive.
Marin is a large-scale pipeline framework built on the Iris job orchestration layer, the Levanter JAX training library, and the Zephyr dataset processing engine. When jobs fail or the controller misbehaves, understanding how to debug issues in Marin requires navigating its distributed architecture and read-only introspection tools. This guide walks through the precise commands and source files you need to diagnose failures without risking state corruption.
Understanding the Debugging Architecture
Marin delegates execution across four distinct layers, each with its own debugging surface. Knowing which component owns your failure is the first step in any troubleshooting workflow.
- Iris (job scheduler): Dispatches jobs, tracks task state in an on-VM SQLite database, and exposes RPC and CLI entry points defined in
lib/iris/src/iris/__init__.py. - Levanter (JAX training): Handles data pipelines, sharding, and checkpointing for large models; reference
lib/levanter/AGENTS.mdfor agent-specific diagnostics. - Zephyr (dataset processing): Manages Parquet reading, writers, and in-memory caching; see
lib/zephyr/OPS.mdfor I/O debugging patterns. - Finelog (time-series metrics): Stores per-worker and per-task metrics separate from the controller DB, accessible via the namespaces defined in
lib/finelog/README.md.
The Iris controller runs as a Docker container and persists its state to $STATE_DIR/db/controller.sqlite3, restoring from GCS checkpoints stored at $REMOTE/controller-state/. According to lib/iris/OPS.md, all Iris interactions are read-only by default, ensuring you can investigate safely before applying mutating commands.
Essential Debugging Commands
Verify Cluster Connectivity
Start every debugging session by confirming the controller is reachable and running the expected image version. Run:
iris --cluster=<NAME> cluster status
This command prints the current Git short-hash and health status. As noted in lib/iris/OPS.md, always verify the image hash matches your intended deployment before proceeding with destructive operations.
Inspect Scheduler State
To view queue length, resource constraints, and priority bands, query the scheduler directly:
iris rpc controller get-scheduler-state
This RPC returns the current allocation state and helps identify bottlenecked resources. For autoscaling decisions, use iris rpc controller get-autoscaler-status and look for the backoff_until_ms field to detect throttling.
Query the Controller Database
The controller stores job and task metadata in SQLite. Use the iris query command to inspect stuck jobs without SSH access:
# Count jobs by state
iris query "SELECT state, count(*) FROM jobs GROUP BY state"
# Export stuck tasks to CSV (state 5 = FAILED)
iris query -f csv "SELECT task_id, state FROM tasks WHERE state=5"
In lib/iris/OPS.md, state 5 corresponds to FAILED, while state 3 indicates RUNNING. Use these queries to identify jobs stuck in BUILDING or RUNNING states before pulling logs.
Analyze Job Logs
Once you identify a suspect job, extract specific error patterns without streaming the entire history:
iris job logs /user/<job-name> --max-lines 400000 --no-tail --substring "Saving checkpoint"
Replace "Saving checkpoint" with "OOM" or "Error" to isolate Out-of-Memory kills or stack traces. The --no-tail flag fetches historical logs rather than following new output.
Profile Running Tasks
When performance degrades, generate CPU or memory flamegraphs without manual SSH or perf installation:
# Generate speedscope JSON for CPU analysis
iris process profile cpu -t /user/job/0
# Generate HTML flamegraph for memory
iris process profile mem -t /user/job/0
These commands, documented in lib/iris/OPS.md, write .speedscope.json or HTML files to your local machine, enabling offline analysis even when the cluster is behind IAP.
Advanced Recovery Techniques
Controller Checkpoint Rollback
If a deployment introduces instability, roll back both the container image and the database state:
iris cluster controller restart --rollback
This restores the previous image and its pre-deployment checkpoint from GCS. According to the source in lib/iris/OPS.md, checkpoints are stored at $REMOTE/controller-state/ and restored to $STATE_DIR/db/controller.sqlite3.
Handling Wedged Controllers
When the controller becomes unresponsive due to database corruption or bloat, never delete the SQLite file. Instead, follow the Checkpoint rollback procedure in lib/iris/OPS.md:
- Move the bloated database aside:
mv $STATE_DIR/db/controller.sqlite3 $STATE_DIR/db/controller.sqlite3.bak - Run
download_checkpoint_to_localinside a temporary container to fetch the last known-good checkpoint from GCS - Restart the controller to load the clean checkpoint
This workflow preserves historical metadata while recovering from disk-full or transaction-lock scenarios.
Automating Debug Workflows with Python
Automate repetitive investigations using the Iris CLI from Python scripts. Save this helper as debug_iris.py:
#!/usr/bin/env python3
import subprocess
import sys
def run(*cmd: str) -> str:
"""Run a shell command and return stdout."""
result = subprocess.run(cmd, capture_output=True, text=True, check=False)
if result.returncode != 0:
print(f"❌ command {' '.join(cmd)} failed:\n{result.stderr}", file=sys.stderr)
sys.exit(1)
return result.stdout.strip()
cluster = "marin"
job = "/user/example-job"
# Verify controller connectivity
print("🔎 Controller status")
print(run("iris", f"--cluster={cluster}", "cluster", "status"))
# Inspect scheduler state
print("\n📊 Scheduler state")
print(run("iris", f"--cluster={cluster}", "rpc", "controller", "get-scheduler-state"))
# Search logs for OOM patterns
print("\n🔎 Searching logs for OOM")
log = run(
"iris", f"--cluster={cluster}", "job", "logs", job,
"--max-lines", "200000", "--substring", "OOM"
)
print(log or "✅ No OOM lines found")
Make the script executable with chmod +x debug_iris.py and invoke it to standardize health checks across your team.
Summary
- Always verify first: Run
iris cluster statusto confirm controller health and image hash before mutating operations. - Query before acting: Use
iris queryto inspect the SQLite database at$STATE_DIR/db/controller.sqlite3to identify stuck jobs without restarting services. - Profile remotely: Generate
.speedscope.jsonflamegraphs viairis process profileto diagnose performance issues without SSH access. - Rollback safely: Use
iris cluster controller restart --rollbackto restore both the container image and database checkpoint when deployments fail. - Never delete the DB: When the controller is wedged, move the SQLite file aside and restore from GCS checkpoints rather than deleting state.
Frequently Asked Questions
How do I check if the Marin controller is healthy?
Run iris --cluster=<NAME> cluster status to verify connectivity and view the current Git short-hash. A healthy controller returns promptly with image metadata; timeouts or connection errors indicate network issues or a crashed container.
What does state=5 mean in Marin job queries?
State 5 represents FAILED jobs in the controller SQLite schema. Use iris query "SELECT job_id, state FROM jobs WHERE state=5" to list failed jobs, then inspect logs with iris job logs /user/<job-name> to retrieve error details.
How can I profile a Marin task without SSH access?
Use the built-in RPC profiler: iris process profile cpu -t /user/job/0 generates a speedscope-compatible JSON file locally. This method works even when workers are behind IAP or in private Kubernetes clusters, unlike manual SSH-based profiling.
How do I safely restart the Marin controller?
First verify the current image hash with iris cluster status. Then run iris cluster controller restart to deploy the current checkout. Never restart without checking the hash, as this command ships the active Git state—deploying from a stale branch will roll back code changes. If the restart causes instability, immediately run iris cluster controller restart --rollback to restore the previous image and checkpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →