How to Troubleshoot Iris Cluster Issues Using OPS.md Diagnostic Patterns

Iris clusters expose a comprehensive set of CLI commands and RPC endpoints that enable read-only inspection of every system layer—from high-level scheduler state to individual process logs and low-level SQLite tables—centralized in lib/iris/OPS.md.

The marin-community/marin repository implements Iris as a distributed job scheduling system with built-in observability tools designed for safe, production debugging. When workloads stall or controllers become unresponsive, operators can follow the structured diagnostic patterns documented in lib/iris/OPS.md to isolate root causes without triggering destructive operations.

The 10-Step Diagnostic Workflow

The official troubleshooting methodology follows a progressive deepening pattern documented in lib/iris/OPS.md, starting with high-level health checks and moving toward low-level database queries and recovery actions. All commands remain read-only unless explicitly marked as mutating.

Step Objective Key Commands OPS.md Reference
1 Identify symptoms iris job list, iris task describe, iris process status Lines 74-82
2 Query scheduler/autoscaler iris rpc controller get-scheduler-state, iris rpc controller get-autoscaler-status Lines 99-104
3 Inspect controller health iris cluster status, iris process status Lines 47-53
4 Analyze logs iris process logs --substring='event=worker_failed' --max-lines 200 Lines 71-78
5 Profile CPU/memory iris process profile cpu -d 10, iris process profile mem Lines 79-84
6 Run SQL queries iris query "SELECT state, count(*) FROM jobs GROUP BY state" Lines 119-128
7 Check rollouts/checkpoints iris cluster controller checkpoint Lines 85-94
8 Roll back deployments iris cluster controller restart --rollback Lines 101-108
9 Execute bulk actions iris query -f csv "$SQL" | iris task preempt --stdin --dry-run Lines 150-166
10 Recover wedged controllers Manual checkpoint restoration procedure Lines 116-129

Phase 1: Symptom Identification and Scheduler Analysis

Start with high-level status checks to determine whether jobs are stuck, workers are missing, or the controller is unresponsive.

To list jobs stuck in PENDING state:

iris job list --state PENDING

If jobs queue but do not progress, inspect the scheduler's internal state to view pending queues and resource constraints:

iris rpc controller get-scheduler-state

Check autoscaler health to identify back-off timers or quota failures preventing worker creation:

iris rpc controller get-autoscaler-status

These RPC endpoints are documented in lib/iris/OPS.md lines 99-104.

Phase 2: Controller Health and Log Analysis

Verify the controller process is running and accessible:

iris cluster status
iris process status

When correlating failures with specific events, filter controller logs using substring matching. This targets the log inspection patterns documented in lines 71-78:

iris process logs --substring='event=worker_failed' --max-lines 200

Phase 3: Performance Profiling

For resource contention or performance anomalies, capture CPU profiles or memory flame-graphs. The following command records a 10-second CPU profile for a specific task:

iris process profile cpu -t /user/job/0 -d 10

Generate memory profiles using:

iris process profile mem

These commands reference the profiling implementations in lib/iris/OPS.md lines 79-84 and leverage helper scripts located at lib/iris/scripts/job_profile_summary.py.

Phase 4: Direct Database Queries

The Iris controller maintains state in SQLite, accessible via read-only SQL queries. This bypasses the RPC layer for raw data inspection as documented in lines 119-128:

iris query "SELECT state, count(*) FROM jobs GROUP BY state"

Query job-specific details or worker assignments using standard SQL syntax against the controller's database schema.

Phase 5: Recovery and Rollback Procedures

Roll Back Bad Deployments

If a new controller image caused cluster instability, restore the previous image and pre-deployment checkpoint using the rollback command documented in lines 101-108:

iris cluster controller restart --rollback

Recover a Wedged Controller (OOM)

When the local database is bloated and causes Out-Of-Memory failures, manually restore a prior checkpoint. This procedure requires stopping the controller container, backing up the corrupted database, and downloading a known-good checkpoint as shown in lines 116-129:

export STATE_DIR=/var/cache/iris/controller
export REMOTE=gs://marin-us-central2/iris/state
sudo docker stop iris-controller
sudo mv "$STATE_DIR/db" "$STATE_DIR/db.bloated.bak.$(date +%s)"
IMAGE=$(sudo docker inspect --format='{{.Config.Image}}' iris-controller)
sudo docker run --rm --network=host -v /var/cache/iris:/var/cache/iris "$IMAGE" \
  .venv/bin/python -c "from pathlib import Path; \
  from iris.cluster.controller.checkpoint import download_checkpoint_to_local as restore; \
  ok = restore('$REMOTE', Path('$STATE_DIR/db'), checkpoint_dir='$REMOTE/controller-state/<epoch_ms>'); \
  raise SystemExit(0 if ok else 1)"
sudo docker start iris-controller

Replace <epoch_ms> with the specific checkpoint timestamp from your remote storage.

Bulk Operations and Automation

For remediation at scale, combine SQL queries with bulk actions. First, identify target tasks using a SQL query, then pipe results into action commands. Always use --dry-run first to preview changes as documented in lines 150-166:

SLICE=marin-tpu-v4-reserved-2048-us-central2-b-...
SQL="SELECT t.task_id FROM tasks t JOIN workers w ON t.current_worker_id=w.worker_id \
     WHERE w.slice_id='$SLICE' AND t.state IN (2,3,9)"
iris query -f csv "$SQL" | iris task preempt --stdin --dry-run

Remove --dry-run to execute the bulk preempt operation.

Key Source Files

Understanding these core files enhances diagnostic capabilities:

Summary

  • lib/iris/OPS.md contains the authoritative 10-step diagnostic pattern for Iris cluster troubleshooting.
  • All diagnostic commands default to read-only operations; mutating actions require explicit flags like --rollback or --stdin.
  • Use iris rpc controller get-scheduler-state to identify scheduling bottlenecks and iris rpc controller get-autoscaler-status for scaling issues.
  • The iris query command provides direct SQL access to the controller's SQLite database for custom state analysis.
  • Recovery from controller OOM or corruption requires manual checkpoint restoration using procedures documented in lines 116-129 of lib/iris/OPS.md.
  • Bulk remediation workflows allow SQL-driven task selection piped into action commands with mandatory --dry-run verification.

Frequently Asked Questions

Where is the Iris troubleshooting documentation located?

The primary diagnostic reference is lib/iris/OPS.md in the marin-community/marin repository. This file contains step-by-step troubleshooting patterns, RPC endpoint documentation, SQL query examples, and recovery procedures ranging from basic health checks to advanced controller resurrection workflows.

How do I check if the Iris scheduler is causing job delays?

Run iris rpc controller get-scheduler-state to view pending queues and resource constraints. If jobs remain in PENDING or BUILDING states, combine this with iris rpc controller get-autoscaler-status to verify whether quota limits or back-off timers are preventing worker provisioning. These commands are documented in lib/iris/OPS.md lines 99-104.

What is the safest way to recover an OOM-wedged Iris controller?

Stop the controller container, archive the bloated database, and restore a prior checkpoint using the Python restoration script referenced in lines 116-129 of lib/iris/OPS.md. This involves downloading a known-good controller.sqlite3.zst from remote storage via download_checkpoint_to_local() before restarting the controller process.

Can I automate bulk actions on stuck Iris tasks?

Yes. Compose a SQL query selecting the target task IDs using iris query -f csv, then pipe the output to iris task preempt, iris task fail, or iris job cancel with the --stdin flag. Always execute with --dry-run first to validate the selection criteria, as documented in lines 150-166 of lib/iris/OPS.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →