How Iris Handles Multi-Cluster Job Scheduling Across GCP and CoreWeave
Iris implements federated scheduling through a three-stage pipeline—classification, placement, and handoff—that routes jobs between GCP-based Marin clusters and CoreWeave GPU clusters based on real-time availability and routing constraints.
Iris, the workload scheduler in the marin-community/marin repository, enables seamless multi-cluster job scheduling across heterogeneous cloud environments. By treating GCP Marin and CoreWeave clusters as peers in a federated topology, Iris allows developers to submit jobs without specifying infrastructure, automatically placing workloads where resources are available while maintaining security boundaries and observability.
Understanding Iris Federated Scheduling Architecture
Iris treats the local Marin cluster and remote CoreWeave clusters as part of a unified federation. Rather than requiring users to manually select infrastructure, the system evaluates job requirements against aggregated resource availability across all peers.
The Three-Stage Scheduling Pipeline
The federated workflow operates through three distinct phases:
- Classification (submit-time) –
PeerRouter.classifyinlib/iris/src/iris/cluster/federation/router.pyevaluates whether a job can run locally, must enter the federation queue, or is unschedulable. - Placement (control tick) –
FederationManager.plan_federationinlib/iris/src/iris/cluster/federation/manager.pyevaluates queued jobs against peer availability reports and promotes candidates to specific backends. - Handoff & Execution – The selected peer receives the job via RPC, authenticates the submitter, and executes the workload using local credentials while the parent cluster retains state mirroring for observability.
Stage 1: Job Classification and Routing Constraints
When a user submits a job, the PeerRouter.classify method applies four deterministic rules to determine placement jurisdiction.
Classification Outcomes
Explicit pin – If the user supplies --target-cluster <peer> (e.g., cw-us-east-02a), the job is immediately queued for that specific CoreWeave peer regardless of local capacity.
Prefer-local – If the local Marin backend can satisfy the job’s resource shape, Iris executes the job on the current cluster without federation overhead.
Peer-queue – When local capacity is insufficient but a reachable peer advertises a backend satisfying all routing constraints—including device-type, device-variant, preemptible, region, and zone—the job enters the federation queue unpinned.
Reject – If no backend in the federation can satisfy the constraints, Iris rejects the job as unschedulable at submission time.
These rules are documented in the federation specification at lib/iris/docs/federation.md.
Stage 2: Placement and Availability Management
Once jobs enter the federation queue, the FederationManager evaluates placement during each scheduler tick using real-time capacity data from all peers.
Availability Heartbeats and Gates
Each peer periodically transmits a heartbeat describing free capacity per backend, including amounts and held_by_band data. During a tick:
- The manager invokes a pure function in
lib/iris/src/iris/cluster/federation/availability.pyto build an availability gate from the job’s device request (e.g., "8 × H100"). - It matches the gate against each peer’s advertised free amount per backend; jobs are never split across backends within a single cluster.
- A promotion is applied via an atomic CAS (compare-and-swap) operation called
promote_queued_handoff, which succeeds only if the job remains queued and uncancelled, preventing race conditions between placement and cancellation.
Priority Band Handling
The placement algorithm respects priority bands when calculating effective capacity. A peer’s available resources are computed as advertised − reserved + capacity held by lower-priority bands, enabling preemptive reclaim of idle capacity from lower-priority workloads.
Stage 3: Cross-Cluster Handoff and Execution
When promotion succeeds, Iris transfers execution authority to the selected peer while preserving job identity and security boundaries.
Authentication and Authorization
The target peer receives the job through the same handoff machinery used for intra-cluster scheduling. The peer:
- Authenticates the upstream Marin cluster using its signed identity.
- Validates the submitter against its
auth.allowed_submitterslist, which receives the forwarded principal from the parent cluster. - Re-creates the job under the same global job ID, ensuring a single job tree exists entirely on the selected cluster.
Credential Isolation
Credentials are not transferred between clouds. Workers in the Marin cluster use GCP service accounts for GCS access, while CoreWeave workers use S3 secrets. Consequently, data artifacts must reside in the destination cluster’s object store before handoff completes.
Operational Examples and CLI Usage
Submit a workload that automatically federates when local H100 capacity is exhausted:
uv run iris job run --device-type=h100 --gpu-count=8 my_training_script.py
The PeerRouter classifies the request as peer-queue eligible, and FederationManager subsequently assigns it to cw-us-east-02a or another available CoreWeave peer.
Force execution on a specific CoreWeave cluster regardless of local availability:
uv run iris job run --target-cluster=cw-us-east-02a \
--device-type=h100 --gpu-count=8 my_training_script.py
This explicit pin causes PeerRouter.classify to return a QUEUE disposition with pinned_peer_id="cw-us-east-02a".
Inspect peer capabilities and real-time capacity:
uv run iris --cluster=marin rpc controller list-peers
Example output shows advertised attributes and availability:
{
"peer_id": "cw-us-east-02a",
"backends": [
{
"advertised_attributes": {"device-type": ["h100"]},
"availability": {"amounts": {"h100": 4}, "held_by_band": []}
}
]
}
Submit programmatically using the Python client:
from iris.client import IrisClient
client = IrisClient(cluster="marin")
job_id = client.submit(
script="my_training_script.py",
device_type="h100",
gpu_count=8,
)
print(f"Submitted job {job_id}")
The client performs identical classification logic to the CLI; the scheduler tick routes the job to CoreWeave if the local backend cannot satisfy the request.
Summary
- Iris federates Marin (GCP) and CoreWeave clusters through a three-stage pipeline involving classification, placement, and handoff.
PeerRouter.classifyinrouter.pydetermines whether jobs run locally, federate, or fail based on routing constraints and explicit pins.FederationManager.plan_federationinmanager.pymatches queued jobs against peer availability using atomic CAS operations to prevent race conditions.- Jobs maintain global identity across clusters while credentials remain isolated to each cloud provider’s native IAM.
- Observability persists through state mirroring, allowing
iris job describeto reflect remote execution status regardless of which cluster hosts the workload.
Frequently Asked Questions
How does Iris determine which cluster handles a job?
Iris evaluates job constraints against real-time availability data from all peers. During the classification phase in lib/iris/src/iris/cluster/federation/router.py, the system checks if the local Marin cluster can satisfy the request. If not, it examines peer heartbeats collected in lib/iris/src/iris/cluster/federation/availability.py to identify a CoreWeave backend with sufficient capacity matching the device-type, region, and zone constraints.
Are job credentials transferred between GCP and CoreWeave?
No. Iris maintains strict credential isolation between cloud providers. When a job hands off to CoreWeave, the Marin cluster does not transfer GCS service account keys. Instead, the CoreWeave cluster executes the job using its own S3 secrets and IAM roles. Users must ensure data artifacts are accessible in the destination cluster’s object store before submission.
What prevents race conditions during job placement?
The FederationManager uses an atomic CAS operation (promote_queued_handoff) when promoting jobs from the federation queue to a specific peer. This operation succeeds only if the job remains in the queued state and has not been cancelled. If the job state changes between evaluation and promotion, the CAS fails and the job remains queued for the next scheduler tick.
How can I monitor jobs that federate to remote clusters?
Iris maintains a state mirror on the parent Marin cluster for all federated jobs. Use iris job list to view federation states including QUEUED_HANDOFF and PENDING_HANDOFF. The iris job describe command aggregates logs from the remote peer through the Finelog relay, providing a unified view of execution status regardless of whether the job runs on GCP or CoreWeave infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →