How to Troubleshoot Iris Cluster Issues Using OPS.md Diagnostic Patterns
Iris clusters expose a comprehensive set of CLI commands and RPC endpoints that enable read-only inspection of every system layer—from high-level scheduler state to individual process logs and low-level SQLite tables—centralized in lib/iris/OPS.md.
The marin-community/marin repository implements Iris as a distributed job scheduling system with built-in observability tools designed for safe, production debugging. When workloads stall or controllers become unresponsive, operators can follow the structured diagnostic patterns documented in lib/iris/OPS.md to isolate root causes without triggering destructive operations.
The 10-Step Diagnostic Workflow
The official troubleshooting methodology follows a progressive deepening pattern documented in lib/iris/OPS.md, starting with high-level health checks and moving toward low-level database queries and recovery actions. All commands remain read-only unless explicitly marked as mutating.
| Step | Objective | Key Commands | OPS.md Reference |
|---|---|---|---|
| 1 | Identify symptoms | iris job list, iris task describe, iris process status |
Lines 74-82 |
| 2 | Query scheduler/autoscaler | iris rpc controller get-scheduler-state, iris rpc controller get-autoscaler-status |
Lines 99-104 |
| 3 | Inspect controller health | iris cluster status, iris process status |
Lines 47-53 |
| 4 | Analyze logs | iris process logs --substring='event=worker_failed' --max-lines 200 |
Lines 71-78 |
| 5 | Profile CPU/memory | iris process profile cpu -d 10, iris process profile mem |
Lines 79-84 |
| 6 | Run SQL queries | iris query "SELECT state, count(*) FROM jobs GROUP BY state" |
Lines 119-128 |
| 7 | Check rollouts/checkpoints | iris cluster controller checkpoint |
Lines 85-94 |
| 8 | Roll back deployments | iris cluster controller restart --rollback |
Lines 101-108 |
| 9 | Execute bulk actions | iris query -f csv "$SQL" | iris task preempt --stdin --dry-run |
Lines 150-166 |
| 10 | Recover wedged controllers | Manual checkpoint restoration procedure | Lines 116-129 |
Phase 1: Symptom Identification and Scheduler Analysis
Start with high-level status checks to determine whether jobs are stuck, workers are missing, or the controller is unresponsive.
To list jobs stuck in PENDING state:
iris job list --state PENDING
If jobs queue but do not progress, inspect the scheduler's internal state to view pending queues and resource constraints:
iris rpc controller get-scheduler-state
Check autoscaler health to identify back-off timers or quota failures preventing worker creation:
iris rpc controller get-autoscaler-status
These RPC endpoints are documented in lib/iris/OPS.md lines 99-104.
Phase 2: Controller Health and Log Analysis
Verify the controller process is running and accessible:
iris cluster status
iris process status
When correlating failures with specific events, filter controller logs using substring matching. This targets the log inspection patterns documented in lines 71-78:
iris process logs --substring='event=worker_failed' --max-lines 200
Phase 3: Performance Profiling
For resource contention or performance anomalies, capture CPU profiles or memory flame-graphs. The following command records a 10-second CPU profile for a specific task:
iris process profile cpu -t /user/job/0 -d 10
Generate memory profiles using:
iris process profile mem
These commands reference the profiling implementations in lib/iris/OPS.md lines 79-84 and leverage helper scripts located at lib/iris/scripts/job_profile_summary.py.
Phase 4: Direct Database Queries
The Iris controller maintains state in SQLite, accessible via read-only SQL queries. This bypasses the RPC layer for raw data inspection as documented in lines 119-128:
iris query "SELECT state, count(*) FROM jobs GROUP BY state"
Query job-specific details or worker assignments using standard SQL syntax against the controller's database schema.
Phase 5: Recovery and Rollback Procedures
Roll Back Bad Deployments
If a new controller image caused cluster instability, restore the previous image and pre-deployment checkpoint using the rollback command documented in lines 101-108:
iris cluster controller restart --rollback
Recover a Wedged Controller (OOM)
When the local database is bloated and causes Out-Of-Memory failures, manually restore a prior checkpoint. This procedure requires stopping the controller container, backing up the corrupted database, and downloading a known-good checkpoint as shown in lines 116-129:
export STATE_DIR=/var/cache/iris/controller
export REMOTE=gs://marin-us-central2/iris/state
sudo docker stop iris-controller
sudo mv "$STATE_DIR/db" "$STATE_DIR/db.bloated.bak.$(date +%s)"
IMAGE=$(sudo docker inspect --format='{{.Config.Image}}' iris-controller)
sudo docker run --rm --network=host -v /var/cache/iris:/var/cache/iris "$IMAGE" \
.venv/bin/python -c "from pathlib import Path; \
from iris.cluster.controller.checkpoint import download_checkpoint_to_local as restore; \
ok = restore('$REMOTE', Path('$STATE_DIR/db'), checkpoint_dir='$REMOTE/controller-state/<epoch_ms>'); \
raise SystemExit(0 if ok else 1)"
sudo docker start iris-controller
Replace <epoch_ms> with the specific checkpoint timestamp from your remote storage.
Bulk Operations and Automation
For remediation at scale, combine SQL queries with bulk actions. First, identify target tasks using a SQL query, then pipe results into action commands. Always use --dry-run first to preview changes as documented in lines 150-166:
SLICE=marin-tpu-v4-reserved-2048-us-central2-b-...
SQL="SELECT t.task_id FROM tasks t JOIN workers w ON t.current_worker_id=w.worker_id \
WHERE w.slice_id='$SLICE' AND t.state IN (2,3,9)"
iris query -f csv "$SQL" | iris task preempt --stdin --dry-run
Remove --dry-run to execute the bulk preempt operation.
Key Source Files
Understanding these core files enhances diagnostic capabilities:
lib/iris/OPS.md— Central reference for all diagnostic commands, RPC calls, and recovery procedures.lib/iris/config/marin.yaml— Default cluster configuration defining endpoints, authentication, and storage backends.lib/iris/scripts/job_profile_summary.py— Helper script for aggregating per-job CPU profiles into flame-graphs.
Summary
lib/iris/OPS.mdcontains the authoritative 10-step diagnostic pattern for Iris cluster troubleshooting.- All diagnostic commands default to read-only operations; mutating actions require explicit flags like
--rollbackor--stdin. - Use
iris rpc controller get-scheduler-stateto identify scheduling bottlenecks andiris rpc controller get-autoscaler-statusfor scaling issues. - The
iris querycommand provides direct SQL access to the controller's SQLite database for custom state analysis. - Recovery from controller OOM or corruption requires manual checkpoint restoration using procedures documented in lines 116-129 of
lib/iris/OPS.md. - Bulk remediation workflows allow SQL-driven task selection piped into action commands with mandatory
--dry-runverification.
Frequently Asked Questions
Where is the Iris troubleshooting documentation located?
The primary diagnostic reference is lib/iris/OPS.md in the marin-community/marin repository. This file contains step-by-step troubleshooting patterns, RPC endpoint documentation, SQL query examples, and recovery procedures ranging from basic health checks to advanced controller resurrection workflows.
How do I check if the Iris scheduler is causing job delays?
Run iris rpc controller get-scheduler-state to view pending queues and resource constraints. If jobs remain in PENDING or BUILDING states, combine this with iris rpc controller get-autoscaler-status to verify whether quota limits or back-off timers are preventing worker provisioning. These commands are documented in lib/iris/OPS.md lines 99-104.
What is the safest way to recover an OOM-wedged Iris controller?
Stop the controller container, archive the bloated database, and restore a prior checkpoint using the Python restoration script referenced in lines 116-129 of lib/iris/OPS.md. This involves downloading a known-good controller.sqlite3.zst from remote storage via download_checkpoint_to_local() before restarting the controller process.
Can I automate bulk actions on stuck Iris tasks?
Yes. Compose a SQL query selecting the target task IDs using iris query -f csv, then pipe the output to iris task preempt, iris task fail, or iris job cancel with the --stdin flag. Always execute with --dry-run first to validate the selection criteria, as documented in lines 150-166 of lib/iris/OPS.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →