# How to Troubleshoot Iris Cluster Issues Using OPS.md Diagnostic Patterns

> Troubleshoot Iris cluster issues with OPS.md diagnostic patterns. Explore CLI commands and RPC endpoints for deep system inspection, from scheduler state to process logs. Learn more now.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Iris clusters expose a comprehensive set of CLI commands and RPC endpoints that enable read-only inspection of every system layer—from high-level scheduler state to individual process logs and low-level SQLite tables—centralized in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md).**

The marin-community/marin repository implements Iris as a distributed job scheduling system with built-in observability tools designed for safe, production debugging. When workloads stall or controllers become unresponsive, operators can follow the structured diagnostic patterns documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) to isolate root causes without triggering destructive operations.

## The 10-Step Diagnostic Workflow

The official troubleshooting methodology follows a progressive deepening pattern documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md), starting with high-level health checks and moving toward low-level database queries and recovery actions. All commands remain **read-only** unless explicitly marked as mutating.

| Step | Objective | Key Commands | OPS.md Reference |
|------|-----------|--------------|------------------|
| 1 | **Identify symptoms** | `iris job list`, `iris task describe`, `iris process status` | Lines 74-82 |
| 2 | **Query scheduler/autoscaler** | `iris rpc controller get-scheduler-state`, `iris rpc controller get-autoscaler-status` | Lines 99-104 |
| 3 | **Inspect controller health** | `iris cluster status`, `iris process status` | Lines 47-53 |
| 4 | **Analyze logs** | `iris process logs --substring='event=worker_failed' --max-lines 200` | Lines 71-78 |
| 5 | **Profile CPU/memory** | `iris process profile cpu -d 10`, `iris process profile mem` | Lines 79-84 |
| 6 | **Run SQL queries** | `iris query "SELECT state, count(*) FROM jobs GROUP BY state"` | Lines 119-128 |
| 7 | **Check rollouts/checkpoints** | `iris cluster controller checkpoint` | Lines 85-94 |
| 8 | **Roll back deployments** | `iris cluster controller restart --rollback` | Lines 101-108 |
| 9 | **Execute bulk actions** | `iris query -f csv "$SQL" \| iris task preempt --stdin --dry-run` | Lines 150-166 |
| 10 | **Recover wedged controllers** | Manual checkpoint restoration procedure | Lines 116-129 |

## Phase 1: Symptom Identification and Scheduler Analysis

Start with high-level status checks to determine whether jobs are stuck, workers are missing, or the controller is unresponsive.

To list jobs stuck in **PENDING** state:

```bash
iris job list --state PENDING

```

If jobs queue but do not progress, inspect the scheduler's internal state to view pending queues and resource constraints:

```bash
iris rpc controller get-scheduler-state

```

Check autoscaler health to identify back-off timers or quota failures preventing worker creation:

```bash
iris rpc controller get-autoscaler-status

```

These RPC endpoints are documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 99-104.

## Phase 2: Controller Health and Log Analysis

Verify the controller process is running and accessible:

```bash
iris cluster status
iris process status

```

When correlating failures with specific events, filter controller logs using substring matching. This targets the log inspection patterns documented in lines 71-78:

```bash
iris process logs --substring='event=worker_failed' --max-lines 200

```

## Phase 3: Performance Profiling

For resource contention or performance anomalies, capture CPU profiles or memory flame-graphs. The following command records a 10-second CPU profile for a specific task:

```bash
iris process profile cpu -t /user/job/0 -d 10

```

Generate memory profiles using:

```bash
iris process profile mem

```

These commands reference the profiling implementations in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 79-84 and leverage helper scripts located at [`lib/iris/scripts/job_profile_summary.py`](https://github.com/marin-community/marin/blob/main/lib/iris/scripts/job_profile_summary.py).

## Phase 4: Direct Database Queries

The Iris controller maintains state in SQLite, accessible via read-only SQL queries. This bypasses the RPC layer for raw data inspection as documented in lines 119-128:

```bash
iris query "SELECT state, count(*) FROM jobs GROUP BY state"

```

Query job-specific details or worker assignments using standard SQL syntax against the controller's database schema.

## Phase 5: Recovery and Rollback Procedures

### Roll Back Bad Deployments

If a new controller image caused cluster instability, restore the previous image and pre-deployment checkpoint using the rollback command documented in lines 101-108:

```bash
iris cluster controller restart --rollback

```

### Recover a Wedged Controller (OOM)

When the local database is bloated and causes Out-Of-Memory failures, manually restore a prior checkpoint. This procedure requires stopping the controller container, backing up the corrupted database, and downloading a known-good checkpoint as shown in lines 116-129:

```bash
export STATE_DIR=/var/cache/iris/controller
export REMOTE=gs://marin-us-central2/iris/state
sudo docker stop iris-controller
sudo mv "$STATE_DIR/db" "$STATE_DIR/db.bloated.bak.$(date +%s)"
IMAGE=$(sudo docker inspect --format='{{.Config.Image}}' iris-controller)
sudo docker run --rm --network=host -v /var/cache/iris:/var/cache/iris "$IMAGE" \
  .venv/bin/python -c "from pathlib import Path; \
  from iris.cluster.controller.checkpoint import download_checkpoint_to_local as restore; \
  ok = restore('$REMOTE', Path('$STATE_DIR/db'), checkpoint_dir='$REMOTE/controller-state/<epoch_ms>'); \
  raise SystemExit(0 if ok else 1)"
sudo docker start iris-controller

```

Replace `<epoch_ms>` with the specific checkpoint timestamp from your remote storage.

## Bulk Operations and Automation

For remediation at scale, combine SQL queries with bulk actions. First, identify target tasks using a SQL query, then pipe results into action commands. Always use `--dry-run` first to preview changes as documented in lines 150-166:

```bash
SLICE=marin-tpu-v4-reserved-2048-us-central2-b-...
SQL="SELECT t.task_id FROM tasks t JOIN workers w ON t.current_worker_id=w.worker_id \
     WHERE w.slice_id='$SLICE' AND t.state IN (2,3,9)"
iris query -f csv "$SQL" | iris task preempt --stdin --dry-run

```

Remove `--dry-run` to execute the bulk preempt operation.

## Key Source Files

Understanding these core files enhances diagnostic capabilities:

- **[`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md)** — Central reference for all diagnostic commands, RPC calls, and recovery procedures.
- **[`lib/iris/config/marin.yaml`](https://github.com/marin-community/marin/blob/main/lib/iris/config/marin.yaml)** — Default cluster configuration defining endpoints, authentication, and storage backends.
- **[`lib/iris/scripts/job_profile_summary.py`](https://github.com/marin-community/marin/blob/main/lib/iris/scripts/job_profile_summary.py)** — Helper script for aggregating per-job CPU profiles into flame-graphs.

## Summary

- **[`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md)** contains the authoritative 10-step diagnostic pattern for Iris cluster troubleshooting.
- All diagnostic commands default to **read-only** operations; mutating actions require explicit flags like `--rollback` or `--stdin`.
- Use `iris rpc controller get-scheduler-state` to identify scheduling bottlenecks and `iris rpc controller get-autoscaler-status` for scaling issues.
- The `iris query` command provides direct SQL access to the controller's SQLite database for custom state analysis.
- Recovery from controller OOM or corruption requires manual checkpoint restoration using procedures documented in lines 116-129 of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md).
- Bulk remediation workflows allow SQL-driven task selection piped into action commands with mandatory `--dry-run` verification.

## Frequently Asked Questions

### Where is the Iris troubleshooting documentation located?

The primary diagnostic reference is **[`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md)** in the marin-community/marin repository. This file contains step-by-step troubleshooting patterns, RPC endpoint documentation, SQL query examples, and recovery procedures ranging from basic health checks to advanced controller resurrection workflows.

### How do I check if the Iris scheduler is causing job delays?

Run `iris rpc controller get-scheduler-state` to view pending queues and resource constraints. If jobs remain in **PENDING** or **BUILDING** states, combine this with `iris rpc controller get-autoscaler-status` to verify whether quota limits or back-off timers are preventing worker provisioning. These commands are documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md) lines 99-104.

### What is the safest way to recover an OOM-wedged Iris controller?

Stop the controller container, archive the bloated database, and restore a prior checkpoint using the Python restoration script referenced in lines 116-129 of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md). This involves downloading a known-good `controller.sqlite3.zst` from remote storage via `download_checkpoint_to_local()` before restarting the controller process.

### Can I automate bulk actions on stuck Iris tasks?

Yes. Compose a SQL query selecting the target task IDs using `iris query -f csv`, then pipe the output to `iris task preempt`, `iris task fail`, or `iris job cancel` with the `--stdin` flag. Always execute with `--dry-run` first to validate the selection criteria, as documented in lines 150-166 of [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md).