How Iris Handles Scheduling Policy and Per‑User Budget Enforcement

Iris enforces scheduling policy and per‑user budget enforcement through a three‑stage pipeline that gates pending tasks, dynamically downgrades priority bands when users exceed spend limits, and round‑robins tasks within bands to prevent monopolization.

The marin-community/marin repository implements Iris as a cluster controller that makes placement decisions every tick by combining hierarchical priority bands with real‑time budget accounting. This article examines the source code to explain how the system builds scheduling contexts, applies budget‑driven band adjustments, and ensures fair resource allocation across competing users.

Building the Scheduling Context

Every controller tick begins with build_scheduling_context in lib/iris/src/iris/cluster/controller/scheduling/policy.py (lines 66‑84). This function aggregates the cluster state required for policy decisions:

  • Healthy workers and their available capacities.
  • Pending tasks mapped to their requested static priority bands.
  • Per‑user spend calculated via compute_user_spend and current budget limits fetched through reads.get_all_user_budget_limits.
  • Default budget limits (UserBudgetDefaults) for users lacking explicit configuration rows.

The resulting context object serves as the immutable input for all downstream scheduling logic, ensuring that every decision uses a consistent snapshot of cluster health and user entitlement.

Gating and Filtering Tasks

Before priority calculation, apply_scheduling_gates (same file, lines 25‑34) filters the pending task set to create GatedCandidates. This step enforces hard constraints such as:

  • Deadline violations — tasks that cannot complete before their deadlines are dropped.
  • Per‑job caps — enforcing max_tasks_per_job_per_cycle to prevent thundering herds.

Only tasks surviving this gate proceed to the priority computation stage, guaranteeing that the scheduler never considers invalid or excessive requests.

Priority Band Calculation with Budget Enforcement

The compute_scheduling_order function (lines 78‑100) translates user spend data into effective scheduling priority through a two‑phase process.

Budget‑Driven Band Downgrade

First, the system builds a task‑to‑band map by calling compute_effective_band from lib/iris/src/iris/cluster/controller/budget.py (lines 67‑88):

task_band_map = {
    task.task_id: compute_effective_band(
        requested_bands.get(task.job_id, job_pb2.PRIORITY_BAND_INTERACTIVE),
        task.task_id.user,
        user_spend,
        user_budget_limits,
        defaults,
    )
}

When a user’s cumulative resource_value spend exceeds their configured budget_limit, compute_effective_band automatically downgrades the task to the BATCH band. This per‑user budget enforcement mechanism ensures that heavy consumers receive lower priority without manual intervention.

System and Production Band Protection

Crucially, the downgrade logic preserves critical infrastructure work. SYSTEM and PRODUCTION bands are never downgraded regardless of spend, maintaining cluster stability and ensuring that high‑priority system‑wide tasks always preempt lower‑priority user workloads.

Fair Scheduling Across Users

Within each priority band, interleave_by_user (lines 91‑112) implements a round‑robin scheduling strategy. The function orders tasks so that users who have consumed less of their budget receive earlier slots. This prevents any single user from monopolizing a priority band while respecting the hierarchical band structure.

Task Assignment and Preemption

After establishing the ordered task list, the controller invokes Scheduler.find_assignments (implemented in lib/iris/src/iris/cluster/controller/scheduling/scheduler.py) to place tasks on workers. This phase respects:

  • Per‑worker limits such as max_building_tasks and max_assignments_per_worker.
  • Device‑variant matching, ensuring that solo pre‑emptors only evict victims with identical accelerator variants.

Preemption Hierarchy

When workers lack capacity, run_preemption_pass (lines 13‑45 in policy.py) enforces a strict eviction hierarchy:

  1. SYSTEM tasks preempt all lower bands.
  2. PRODUCTION tasks preempt INTERACTIVE and BATCH.
  3. INTERACTIVE tasks preempt BATCH only.
  4. BATCH tasks never preempt others.

Preemptions are only permitted across different priority bands; same‑band tasks compete exclusively through the ordering established earlier.

Transactional Persistence

Once assignments and preemptions are selected, the controller persists decisions in a single database transaction using writes.tasks.assign_task and writes.tasks.preempt_task. This atomic write guarantees that a pre‑emptor never re‑competes for the slot it just freed, eliminating scheduling races.

Summary

  • Iris implements scheduling policy and per‑user budget enforcement through a declarative three‑stage pipeline defined in policy.py.
  • Budget limits trigger automatic downgrades to the BATCH band via compute_effective_band in budget.py, while protecting SYSTEM and PRODUCTION workloads.
  • Fairness is guaranteed by round‑robin interleaving within bands, preventing monopolization by heavy consumers.
  • Preemption follows a strict hierarchy that respects band priority, and all decisions are committed atomically to maintain consistency.

Frequently Asked Questions

How does Iris prevent a single user from consuming all cluster resources?

Within each priority band, interleave_by_user round‑robins tasks based on current spend, ensuring that users with lower utilization get earlier scheduling slots. Additionally, when a user exceeds their budget_limit, all their tasks downgrade to the BATCH band, where they compete with other deprioritized workloads and can be preempted by higher‑priority bands.

What happens to SYSTEM and PRODUCTION tasks when a user exceeds their budget?

SYSTEM and PRODUCTION bands are never downgraded regardless of spend. The compute_effective_band function explicitly preserves these bands to maintain cluster infrastructure stability and ensure that critical system workloads retain highest priority even when individual users exhaust their allocations.

How does Iris handle task deadlines during scheduling?

The apply_scheduling_gates function filters out tasks that cannot meet their deadlines before they enter the priority calculation phase. This gating happens early in the pipeline (lines 25‑34 of policy.py), ensuring that the scheduler never wastes cycles on infeasible assignments and that deadline‑sensitive jobs fail fast rather than queue indefinitely.

Where can I find tests validating the budget enforcement logic?

The repository includes comprehensive validation in lib/iris/tests/test_budget.py, which covers spend tracking and band downgrade behavior. Scheduling fairness across priority bands is verified in lib/iris/tests/cluster/controller/test_scheduling_fairness.py, ensuring that the hierarchy and interleaving logic function correctly under diverse load patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →