# How Iris Autoscaler Worker Provisioning and Scaling Works: A Code-Level Breakdown

> Explore the Iris autoscaler's code-level logic for worker provisioning and scaling. Understand how it manages cluster state, job demand, capacity, and failure handling.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: internals
- Published: 2026-08-29

---

**The Iris autoscaler automatically provisions and scales workers by running a periodic refresh cycle that collects cluster state, computes demand from pending jobs, routes entries to scaling groups, evaluates capacity and quota constraints, and issues platform-specific provisioning requests—while handling failures through classified back-off markers and stock-out detection.**

The `marin-community/marin` repository contains Iris, a distributed job scheduler with a sophisticated autoscaler designed for heterogeneous compute environments. The autoscaler lives in `iris/cluster/controller/autoscaler/` and implements deterministic worker lifecycle management through state-driven decision pipelines.

## Core Autoscaler Architecture

The autoscaler operates as a **state machine** that transforms cluster observations into provisioning actions. At its heart lies the `Status` object ([`iris/cluster/controller/autoscaler/status.py`](https://github.com/marin-community/marin/blob/main/iris/cluster/controller/autoscaler/status.py)), which aggregates platform state, quota consumption, and historical failure markers before each decision cycle.

### The Refresh Cycle Pipeline

Each autoscaler iteration follows a strict seven-step pipeline:

1. **State Collection** – Query platform APIs for live worker inventory
2. **Demand Computation** – Convert pending jobs into typed `DemandEntry` objects
3. **Demand Routing** – Assign entries to compatible `ScalingGroup` instances
4. **Group Evaluation** – Check capacity, quota, back-off, and stock-out conditions
5. **Provisioning Orchestration** – Execute platform-specific create calls
6. **Failure Classification** – Map failures to retry policies via `classify_create_failure()`
7. **Status Emission** – Publish actions and `PendingHint` objects for downstream consumers

This pipeline executes atomically: either all groups are evaluated or the cycle aborts, preventing partial provisioning states that could strand resources.

## Worker Provisioning Logic

Provisioning decisions emerge from the interaction between three core abstractions: **demand entries**, **scaling groups**, and **group availability states**.

### DemandEntry and Resource Modeling

The autoscaler represents compute requirements through `DemandEntry` instances defined in [`iris/cluster/controller/autoscaler/models.py`](https://github.com/marin-community/marin/blob/main/iris/cluster/controller/autoscaler/models.py). These encapsulate:

- Resource requests (CPU, GPU type/count, TPU topology)
- Placement constraints (zones, machine families, spot/preemptible preferences)
- Job priority and scheduling class

```python
from iris.cluster.controller.autoscaler.models import DemandEntry

# A TPU v5-8 job requesting 2 workers with zone affinity

demand = DemandEntry(
    resources={"tpu": {"version": "v5", "topology": "2x4", "count": 2}},
    constraints=["us-central1-b", "preemptible"],
    priority=100
)

```

### ScalingGroup Evaluation

Each `ScalingGroup` ([`iris/cluster/controller/autoscaler/scaling_group.py`](https://github.com/marin-community/marin/blob/main/iris/cluster/controller/autoscaler/scaling_group.py)) maintains independent capacity accounting. The autoscaler evaluates groups through a priority-ordered filter chain:

| Check | Failure Outcome | Metric Surfaced |
|-------|---------------|---------------|
| Capacity saturation | Skip group, mark `CAPACITY_EXCEEDED` | `group.capacity_remaining` |
| Quota exhaustion | Skip group, mark `QUOTA_EXCEEDED` | `group.quota_remaining` |
| Active back-off window | Skip group, increment `backoff_skipped` counter | `group.backoff_expires_at` |
| Stock-out marker | Skip group, log `STOCKOUT_MARKER` | `group.stockout_until` |

Only groups passing all filters proceed to provisioning.

### Stock-Out and Back-Off Handling

The autoscaler implements **adaptive retry policies** through `classify_create_failure()` in [`iris/cluster/controller/autoscaler/provisioning.py`](https://github.com/marin-community/marin/blob/main/iris/cluster/controller/autoscaler/provisioning.py). Platform errors map to durable state:

```python
from iris.cluster.controller.autoscaler.provisioning import classify_create_failure

# Example: GPU capacity exhaustion on CoreWeave

error = "RESOURCE_EXHAUSTED: nvidia-tesla-a100 not available in us-east1"
classification = classify_create_failure(error)

# Returns: "stockout" → triggers 15-minute stock-out marker

# Alternative returns: "transient" (30s backoff), "quota" (quota-specific backoff), "permanent" (no retry)

```

Stock-out markers persist across refresh cycles, preventing thundering-herd retry storms against depleted resource pools.

## Demand Routing and Group Selection

The routing layer ([`iris/cluster/controller/autoscaler/routing.py`](https://github.com/marin-community/marin/blob/main/iris/cluster/controller/autoscaler/routing.py)) implements **feasibility-based assignment** using job requirements and group capabilities. Routing respects:

- **Device type monotonicity** – Jobs requesting specific accelerator versions route only to compatible groups
- **Tier affinity** – Production-tier jobs avoid spot/preemptible groups unless explicitly allowed
- **Quota-pool isolation** – Enforced through `quota-pool tier monotonicity` constraints that prevent cross-pool contamination

When no viable group exists, the autoscaler generates `PendingHint` objects explaining the blockage:

```python
from iris.cluster.controller.autoscaler.status import build_job_pending_hints

# After a failed routing pass

hints = build_job_pending_hints(autoscaler.get_status())
for hint in hints:
    print(f"{hint.job_id}: {hint.reason}")  # e.g., "tier_blocked: no production group available"

```

These hints surface in logs as structured unsatisfied demand records:

```

Unsatisfied autoscaler demand: tier_blocked: job=abc123, requested=TPU-v5-8, available_groups=0

```

## Scaling Decisions: Up and Down

### Scale-Up Triggering

Scale-up occurs when **routed demand exceeds satisfied capacity** after all group evaluations. The autoscaler:

1. Calculates **deficit** per group: `max(0, assigned_demand - current_capacity - pending_provisions)`
2. Batches provisioning requests by **slice size** – atomic units of worker allocation
3. Submits platform API calls through abstracted provisioners (CoreWeave, GCP, local VM)

```python
from iris.testing.controller import make_autoscaler
from iris.testing.platform import make_mock_platform
from iris.cluster.controller.autoscaler.scaling_group import ScalingGroup

# Test setup: single TPU group with 8-node capacity

platform = make_mock_platform()
group = ScalingGroup({"name": "tpu-v5-8", "max_workers": 8}, platform)
autoscaler = make_autoscaler({"tpu-v5-8": group})

# Simulate demand for 12 workers (4 over capacity)

autoscaler.run_once(
    demand=[DemandEntry(resources={"tpu": 12}, constraints=["v5-8"])],
    status={},
    timestamp=Timestamp.now()
)

# Result: scale-up action for 8 workers, 4 unsatisfied with tier_blocked hint

```

### Scale-Down and Idle Reclamation

Worker termination follows **conservative idle detection**: a group must report zero active tasks for `idle_threshold` duration (default 300s) before scale-down eligibility. The autoscaler:

- Respects **drain-group policies** that extend idle windows for graceful job hand-off
- Terminates workers in **slice granularity** to maintain cluster topology invariants
- Updates `Status.recent_actions` with `scale_down` entries for observability

## Integration Testing and Monitoring Paths

The autoscaler's behavior is validated through comprehensive test suites:

| Test File | Coverage |
|-----------|----------|
| [`lib/iris/tests/cluster/controller/test_provisioning.py`](https://github.com/marin-community/marin/blob/main/lib/iris/tests/cluster/controller/test_provisioning.py) | Unit tests for provisioning flow, failure classification, stock-out markers |
| [`tests/ci/test_iris_monitor.py`](https://github.com/marin-community/marin/blob/main/tests/ci/test_iris_monitor.py) | Integration tests verifying unsatisfied demand logging and hint generation |

These tests use `make_autoscaler()` from [`iris/testing/controller.py`](https://github.com/marin-community/marin/blob/main/iris/testing/controller.py) to inject mock platforms and deterministic state transitions.

## Summary

- **The Iris autoscaler implements a state-driven pipeline** – collect, route, evaluate, provision, emit – executed atomically per refresh cycle
- **Provisioning decisions respect four guardrails** – capacity, quota, back-off windows, and stock-out markers – with failure classification via `classify_create_failure()`
- **Demand routing uses feasibility filters** including device type monotonicity and tier affinity, surfacing blockages through `PendingHint` objects
- **Scale-up triggers on unsatisfied routed demand**; scale-down requires idle threshold satisfaction with drain-group policy compliance
- **Core implementation spans** [`provisioning.py`](https://github.com/marin-community/marin/blob/main/provisioning.py), [`scaling_group.py`](https://github.com/marin-community/marin/blob/main/scaling_group.py), [`routing.py`](https://github.com/marin-community/marin/blob/main/routing.py), [`models.py`](https://github.com/marin-community/marin/blob/main/models.py), and [`status.py`](https://github.com/marin-community/marin/blob/main/status.py) under `iris/cluster/controller/autoscaler/`

## Frequently Asked Questions

### How does the Iris autoscaler handle platform capacity exhaustion?

The autoscaler detects capacity exhaustion through `classify_create_failure()`, which maps platform errors like `RESOURCE_EXHAUSTED` to a `STOCKOUT_MARKER`. This marker persists for 15 minutes, causing the affected `ScalingGroup` to be skipped in subsequent routing passes. The marker prevents wasted API calls and allows demand to flow to alternative groups or regions if configured.

### What determines which scaling group receives a pending job?

The routing layer ([`routing.py`](https://github.com/marin-community/marin/blob/main/routing.py)) assigns `DemandEntry` objects to groups using feasibility filtering: device type compatibility, zone constraints, tier requirements (production vs. spot), and quota-pool membership. When multiple groups qualify, the autoscaler preferentially selects groups with higher remaining capacity and no active back-off markers.

### Why would the autoscaler log "tier_blocked" for unsatisfied demand?

A `tier_blocked` hint appears when a job's scheduling tier (e.g., production) cannot be satisfied by any available group, typically because all compatible groups are preemptible/spot instances or lack quota in the production pool. This enforces quota-pool tier monotonicity – production work never runs on spot infrastructure unless explicitly permitted by policy.

### How can I test autoscaler behavior without a live cluster?

Use the testing utilities in [`iris/testing/controller.py`](https://github.com/marin-community/marin/blob/main/iris/testing/controller.py). The `make_autoscaler()` factory accepts mock platforms from `make_mock_platform()` and allows injection of arbitrary `ScalingGroup` configurations. The test suite in [`lib/iris/tests/cluster/controller/test_provisioning.py`](https://github.com/marin-community/marin/blob/main/lib/iris/tests/cluster/controller/test_provisioning.py) demonstrates deterministic state transitions, failure injection, and assertion patterns for custom autoscaler scenarios.