WORKER vs KUBERNETES Backend Kinds in Iris: Execution Models Explained

The WORKER backend runs the Iris scheduler locally and manages workers via RPC, while the KUBERNETES backend delegates placement to Kueue and reconciles Pods directly.

Iris, the distributed task orchestrator in marin-community/marin, supports two distinct backend execution models through its BackendKind enum. Understanding the differences between WORKER and KUBERNETES backend kinds is critical for choosing the right deployment architecture for your cluster. This guide breaks down the technical distinctions, code paths, and use cases for each backend type.


BackendKind Enum Definition

The BackendKind enum in iris/cluster/controller/backend.py provides the foundational distinction:

from enum import StrEnum

class BackendKind(StrEnum):
    """The execution mechanism owned by a controller's backend."""
    WORKER = "worker‑daemon"
    KUBERNETES = "kubernetes"

Each BackendDescriptor carries this kind field, which determines the controller's entire runtime behavior. The backend descriptor also includes backend_id, advertised_attributes, scale_groups, and display_name, but the kind field is the primary selector for conditional logic throughout the controller stack.


Execution Model: Local Scheduler vs. Kubernetes Delegation

WORKER Backend Runs Iris Scheduler Locally

A WORKER backend operates in Iris-centric mode. The controller runs the native Iris scheduler and communicates with workers directly through RPC calls. This model gives Iris full control over task placement, worker health, and autoscaling.

Key characteristics of the WORKER backend:

  • Scheduler location: Runs inside the controller process
  • Worker communication: Direct RPC via worker network addresses
  • State ownership: Maintains a ControllerDB worker store tracking active workers, health status, and resource availability
  • Autoscaling: Uses WorkerHealthTracker for worker-specific reconcile and autoscaling logic

In iris/cluster/controller/controller.py, worker-only actions are guarded with explicit kind checks:

if descriptor.kind is BackendKind.WORKER:
    # Health tracking, autoscaling, and worker store updates

    self._update_worker_health()
    self._trigger_autoscale()

KUBERNETES Backend Delegates to Kueue

A KUBERNETES backend operates in substrate-delegated mode. Instead of running its own scheduler, the controller hands placement decisions to Kueue, the Kubernetes-native job queueing system. The controller's role shifts to Pod lifecycle management.

Key characteristics of the KUBERNETES backend:

  • Scheduler location: External (Kueue manages queueing and placement)
  • Worker communication: Uses attempt_uid to rebuild pod names; no network addresses needed
  • State ownership: No worker store — placement state lives in the Kubernetes control plane
  • Autoscaling: Relies on cluster-level autoscaling (Karpenter, Cluster Autoscaler, etc.)

The controller skips worker-specific reconcile logic for Kubernetes backends, as seen in iris/cluster/controller/controller.py around lines 1040-1052:

if descriptor.kind is BackendKind.KUBERNETES:
    # Direct Pod reconciliation, no worker health tracking

    self._reconcile_pods()

State Handling: Worker Store vs. Cluster Substrate

The presence or absence of a worker store is one of the most significant architectural differences between the two backend kinds.

Component WORKER KUBERNETES
Worker store (ControllerDB) ✅ Required ❌ Absent
Health tracking WorkerHealthTracker class Kubernetes readiness probes
Resource accounting Manual via worker store Kueue quotas and flavors
Worker lifecycle Explicit join/leave RPC Pod creation/deletion events

For WORKER backends, the ControllerDB maintains active worker records with their network addresses, capabilities, and health status. This enables fine-grained scheduling decisions and rapid failure detection.

For KUBERNETES backends, the controller trusts the cluster substrate. Worker identity maps to Pod identity through attempt_uid-derived names, and health is determined by Kubernetes-native mechanisms rather than custom tracking.


Dashboard Capabilities: What Metrics Each Backend Exposes

The dashboard_backend_descriptor function in iris/cluster/controller/backend.py selects capabilities based on backend kind:

def dashboard_backend_descriptor(controller):
    desc = controller.backend_descriptor
    if desc.kind is BackendKind.WORKER:
        return DashboardDescriptor(
            name=desc.display_name,
            capabilities=["workers", "autoscaler"]  # if autoscaling configured

        )
    else:  # KUBERNETES

        return DashboardDescriptor(
            name=desc.display_name,
            capabilities=["cluster"]
        )

This means:

  • WORKER backends expose worker-specific metrics and autoscaling status
  • KUBERNETES backends expose only cluster-level metrics, delegating detailed visibility to Kubernetes tooling (kubectl, Kueue dashboards, etc.)

Task Targeting: Address-Based vs. UID-Based RPC Routing

Task RPC addressing differs fundamentally between the two backends:

WORKER: Network Address Targeting


# TaskTarget contains a direct network address

task_target = TaskTarget(address="10.0.0.15:9090")

# RPC routed to specific worker daemon

KUBERNETES: attempt_uid to Pod Name Resolution


# TaskTarget carries attempt_uid, controller rebuilds pod name

task_target = TaskTarget(attempt_uid="task-abc-123-xyz")

# Controller constructs: f"{task_name}-{attempt_uid}-{container_name}"

# No static address needed — Kubernetes DNS resolves the Pod

This design lets Kubernetes backends work with ephemeral Pod identities while maintaining stable task references through attempt_uid.


Practical Examples: Configuring Each Backend Kind

Creating a WORKER Backend Descriptor

from iris.cluster.controller.backend import BackendDescriptor, BackendKind

worker_backend = BackendDescriptor(
    backend_id="local‑worker‑01",
    kind=BackendKind.WORKER,
    advertised_attributes={"gpu": {"v100"}},
    scale_groups=frozenset({"default"}),
    display_name="Local Worker Backend",
)

# Controller behavior:

#   * Native Iris scheduler runs locally

#   * RPC to workers via stored network addresses

#   * Worker health tracked via WorkerHealthTracker

#   * Dashboard shows "workers" + "autoscaler" tabs

Creating a KUBERNETES Backend Descriptor

from iris.cluster.controller.backend import BackendDescriptor, BackendKind

k8s_backend = BackendDescriptor(
    backend_id="k8s‑cluster‑01",
    kind=BackendKind.KUBERNETES,
    advertised_attributes={"gpu": {"a100"}},
    scale_groups=frozenset({"high‑mem"}),
    display_name="K8s Cluster Backend",
)

# Controller behavior:

#   * Placement delegated to Kueue

#   * Pods reconciled directly via Kubernetes API

#   * No worker store maintained

#   * Dashboard shows "cluster" tab only

Runtime Backend Inspection

from iris.cluster.controller.backend import dashboard_backend_descriptor

def print_capabilities(controller):
    desc = dashboard_backend_descriptor(controller)
    print(f"Backend '{desc.name}' supports: {', '.join(desc.capabilities)}")

# WORKER output: "Backend 'Local Worker Backend' supports: workers, autoscaler"

# KUBERNETES output: "Backend 'K8s Cluster Backend' supports: cluster"

When to Use Each Backend Kind

Scenario Recommended Backend Rationale
Small-to-medium cluster (<100 nodes) WORKER Direct control, lower operational complexity, no Kubernetes required
Large Kubernetes fleet KUBERNETES Leverages existing Kueue infrastructure, native autoscaling, multi-tenant queues
Custom hardware schedulers WORKER Native scheduler can be extended with custom placement policies
Serverless/bursty workloads KUBERNETES Kueue + Karpenter handles rapid scale-out without controller bottlenecks
Hybrid cloud KUBERNETES Consistent abstraction across cloud providers

Key Source Files and Their Roles

File Path Purpose
iris/cluster/controller/backend.py Defines BackendKind, BackendDescriptor, dashboard_backend_descriptor()
iris/cluster/controller/controller.py Main controller loop with kind-based branching (lines 540-588 for WORKER, 1040-1052 for KUBERNETES)
iris/cluster/controller/service.py RPC routing logic that adapts to backend kind
iris/tests/journeys/backend.py Integration tests validating both backend behaviors
iris/tests/cluster/controller/test_direct_controller.py Unit tests for conditional backend paths

The module docstring in iris/cluster/controller/backend.py (lines 18-22) explicitly documents this architectural split:

"The two backend kinds have deliberately different mechanisms: a worker backend runs the Iris scheduler and worker RPCs, while a Kubernetes backend hands placement to Kueue and reconciles Pods directly."


Summary

  • WORKER backends run the native Iris scheduler, maintain a worker store, track health via WorkerHealthTracker, and expose worker/autoscaler dashboard capabilities
  • KUBERNETES backends delegate scheduling to Kueue, have no worker store, reconcile Pods directly, and expose only cluster-level dashboard metrics
  • The BackendKind enum in iris/cluster/controller/backend.py drives conditional logic throughout the controller stack
  • Task targeting uses network addresses for WORKER and attempt_uid-based pod names for KUBERNETES
  • Choose WORKER for direct control and smaller deployments; choose KUBERNETES for large-scale Kubernetes-native operation

Frequently Asked Questions

Can a single Iris controller switch between WORKER and KUBERNETES backends at runtime?

No. The BackendDescriptor is set at controller initialization and immutable thereafter. The controller's entire execution path—including scheduler instantiation, state management, and RPC handling—is determined by the kind field. To change backends, you must restart the controller with a new descriptor.

Does the KUBERNETES backend support custom Iris scheduling policies?

No. Custom scheduling policies require the native Iris scheduler, which only runs with BackendKind.WORKER. The KUBERNETES backend intentionally relinquishes scheduling control to Kueue, inheriting its queueing and placement semantics. For custom policies in Kubernetes environments, consider extending Kueue's flavors or using WORKER backends with a Kubernetes-sidecar deployment model.

How does autoscaling differ between the two backend kinds?

WORKER backends use Iris-native autoscaling via WorkerHealthTracker and explicit scale-group logic in the controller. KUBERNETES backends rely entirely on cluster-level autoscaling: Karpenter, Cluster Autoscaler, or node pools responding to pending Pod conditions. The controller does not perform worker-specific autoscale calculations for Kubernetes backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →