How Iris Autoscaler Worker Provisioning and Scaling Works: A Code-Level Breakdown
The Iris autoscaler automatically provisions and scales workers by running a periodic refresh cycle that collects cluster state, computes demand from pending jobs, routes entries to scaling groups, evaluates capacity and quota constraints, and issues platform-specific provisioning requests—while handling failures through classified back-off markers and stock-out detection.
The marin-community/marin repository contains Iris, a distributed job scheduler with a sophisticated autoscaler designed for heterogeneous compute environments. The autoscaler lives in iris/cluster/controller/autoscaler/ and implements deterministic worker lifecycle management through state-driven decision pipelines.
Core Autoscaler Architecture
The autoscaler operates as a state machine that transforms cluster observations into provisioning actions. At its heart lies the Status object (iris/cluster/controller/autoscaler/status.py), which aggregates platform state, quota consumption, and historical failure markers before each decision cycle.
The Refresh Cycle Pipeline
Each autoscaler iteration follows a strict seven-step pipeline:
- State Collection – Query platform APIs for live worker inventory
- Demand Computation – Convert pending jobs into typed
DemandEntryobjects - Demand Routing – Assign entries to compatible
ScalingGroupinstances - Group Evaluation – Check capacity, quota, back-off, and stock-out conditions
- Provisioning Orchestration – Execute platform-specific create calls
- Failure Classification – Map failures to retry policies via
classify_create_failure() - Status Emission – Publish actions and
PendingHintobjects for downstream consumers
This pipeline executes atomically: either all groups are evaluated or the cycle aborts, preventing partial provisioning states that could strand resources.
Worker Provisioning Logic
Provisioning decisions emerge from the interaction between three core abstractions: demand entries, scaling groups, and group availability states.
DemandEntry and Resource Modeling
The autoscaler represents compute requirements through DemandEntry instances defined in iris/cluster/controller/autoscaler/models.py. These encapsulate:
- Resource requests (CPU, GPU type/count, TPU topology)
- Placement constraints (zones, machine families, spot/preemptible preferences)
- Job priority and scheduling class
from iris.cluster.controller.autoscaler.models import DemandEntry
# A TPU v5-8 job requesting 2 workers with zone affinity
demand = DemandEntry(
resources={"tpu": {"version": "v5", "topology": "2x4", "count": 2}},
constraints=["us-central1-b", "preemptible"],
priority=100
)
ScalingGroup Evaluation
Each ScalingGroup (iris/cluster/controller/autoscaler/scaling_group.py) maintains independent capacity accounting. The autoscaler evaluates groups through a priority-ordered filter chain:
| Check | Failure Outcome | Metric Surfaced |
|---|---|---|
| Capacity saturation | Skip group, mark CAPACITY_EXCEEDED |
group.capacity_remaining |
| Quota exhaustion | Skip group, mark QUOTA_EXCEEDED |
group.quota_remaining |
| Active back-off window | Skip group, increment backoff_skipped counter |
group.backoff_expires_at |
| Stock-out marker | Skip group, log STOCKOUT_MARKER |
group.stockout_until |
Only groups passing all filters proceed to provisioning.
Stock-Out and Back-Off Handling
The autoscaler implements adaptive retry policies through classify_create_failure() in iris/cluster/controller/autoscaler/provisioning.py. Platform errors map to durable state:
from iris.cluster.controller.autoscaler.provisioning import classify_create_failure
# Example: GPU capacity exhaustion on CoreWeave
error = "RESOURCE_EXHAUSTED: nvidia-tesla-a100 not available in us-east1"
classification = classify_create_failure(error)
# Returns: "stockout" → triggers 15-minute stock-out marker
# Alternative returns: "transient" (30s backoff), "quota" (quota-specific backoff), "permanent" (no retry)
Stock-out markers persist across refresh cycles, preventing thundering-herd retry storms against depleted resource pools.
Demand Routing and Group Selection
The routing layer (iris/cluster/controller/autoscaler/routing.py) implements feasibility-based assignment using job requirements and group capabilities. Routing respects:
- Device type monotonicity – Jobs requesting specific accelerator versions route only to compatible groups
- Tier affinity – Production-tier jobs avoid spot/preemptible groups unless explicitly allowed
- Quota-pool isolation – Enforced through
quota-pool tier monotonicityconstraints that prevent cross-pool contamination
When no viable group exists, the autoscaler generates PendingHint objects explaining the blockage:
from iris.cluster.controller.autoscaler.status import build_job_pending_hints
# After a failed routing pass
hints = build_job_pending_hints(autoscaler.get_status())
for hint in hints:
print(f"{hint.job_id}: {hint.reason}") # e.g., "tier_blocked: no production group available"
These hints surface in logs as structured unsatisfied demand records:
Unsatisfied autoscaler demand: tier_blocked: job=abc123, requested=TPU-v5-8, available_groups=0
Scaling Decisions: Up and Down
Scale-Up Triggering
Scale-up occurs when routed demand exceeds satisfied capacity after all group evaluations. The autoscaler:
- Calculates deficit per group:
max(0, assigned_demand - current_capacity - pending_provisions) - Batches provisioning requests by slice size – atomic units of worker allocation
- Submits platform API calls through abstracted provisioners (CoreWeave, GCP, local VM)
from iris.testing.controller import make_autoscaler
from iris.testing.platform import make_mock_platform
from iris.cluster.controller.autoscaler.scaling_group import ScalingGroup
# Test setup: single TPU group with 8-node capacity
platform = make_mock_platform()
group = ScalingGroup({"name": "tpu-v5-8", "max_workers": 8}, platform)
autoscaler = make_autoscaler({"tpu-v5-8": group})
# Simulate demand for 12 workers (4 over capacity)
autoscaler.run_once(
demand=[DemandEntry(resources={"tpu": 12}, constraints=["v5-8"])],
status={},
timestamp=Timestamp.now()
)
# Result: scale-up action for 8 workers, 4 unsatisfied with tier_blocked hint
Scale-Down and Idle Reclamation
Worker termination follows conservative idle detection: a group must report zero active tasks for idle_threshold duration (default 300s) before scale-down eligibility. The autoscaler:
- Respects drain-group policies that extend idle windows for graceful job hand-off
- Terminates workers in slice granularity to maintain cluster topology invariants
- Updates
Status.recent_actionswithscale_downentries for observability
Integration Testing and Monitoring Paths
The autoscaler's behavior is validated through comprehensive test suites:
| Test File | Coverage |
|---|---|
lib/iris/tests/cluster/controller/test_provisioning.py |
Unit tests for provisioning flow, failure classification, stock-out markers |
tests/ci/test_iris_monitor.py |
Integration tests verifying unsatisfied demand logging and hint generation |
These tests use make_autoscaler() from iris/testing/controller.py to inject mock platforms and deterministic state transitions.
Summary
- The Iris autoscaler implements a state-driven pipeline – collect, route, evaluate, provision, emit – executed atomically per refresh cycle
- Provisioning decisions respect four guardrails – capacity, quota, back-off windows, and stock-out markers – with failure classification via
classify_create_failure() - Demand routing uses feasibility filters including device type monotonicity and tier affinity, surfacing blockages through
PendingHintobjects - Scale-up triggers on unsatisfied routed demand; scale-down requires idle threshold satisfaction with drain-group policy compliance
- Core implementation spans
provisioning.py,scaling_group.py,routing.py,models.py, andstatus.pyunderiris/cluster/controller/autoscaler/
Frequently Asked Questions
How does the Iris autoscaler handle platform capacity exhaustion?
The autoscaler detects capacity exhaustion through classify_create_failure(), which maps platform errors like RESOURCE_EXHAUSTED to a STOCKOUT_MARKER. This marker persists for 15 minutes, causing the affected ScalingGroup to be skipped in subsequent routing passes. The marker prevents wasted API calls and allows demand to flow to alternative groups or regions if configured.
What determines which scaling group receives a pending job?
The routing layer (routing.py) assigns DemandEntry objects to groups using feasibility filtering: device type compatibility, zone constraints, tier requirements (production vs. spot), and quota-pool membership. When multiple groups qualify, the autoscaler preferentially selects groups with higher remaining capacity and no active back-off markers.
Why would the autoscaler log "tier_blocked" for unsatisfied demand?
A tier_blocked hint appears when a job's scheduling tier (e.g., production) cannot be satisfied by any available group, typically because all compatible groups are preemptible/spot instances or lack quota in the production pool. This enforces quota-pool tier monotonicity – production work never runs on spot infrastructure unless explicitly permitted by policy.
How can I test autoscaler behavior without a live cluster?
Use the testing utilities in iris/testing/controller.py. The make_autoscaler() factory accepts mock platforms from make_mock_platform() and allows injection of arbitrary ScalingGroup configurations. The test suite in lib/iris/tests/cluster/controller/test_provisioning.py demonstrates deterministic state transitions, failure injection, and assertion patterns for custom autoscaler scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →