How to Add a New Cluster Backend Provider to Iris: Complete Implementation Guide

To add a new cluster backend provider to Iris, implement the TaskBackend protocol in a new module under lib/iris/src/iris/cluster/backends/, register the class in the BACKEND_REGISTRY dictionary in lib/iris/src/iris/cluster/backends/__init__.py, and update the CLI parser in lib/iris/src/iris/cli.py to recognize the new provider identifier.

Iris, the distributed task orchestrator in the marin-community/marin repository, uses a pluggable backend architecture to abstract cluster operations. When you add a new cluster backend provider to Iris, you enable the controller to schedule, reconcile, and autoscale workloads on your custom infrastructure. This guide walks through the protocol implementation, registration process, and integration steps required to extend Iris with a new backend.

Understand the TaskBackend Protocol

The TaskBackend protocol defines the contract between Iris and cluster infrastructure. Defined in lib/iris/src/iris/cluster/controller/backend.py, it requires 15 concrete implementations that handle the complete lifecycle of task scheduling, execution, and resource management.

Your implementation must provide these core operational methods:

  • Scheduling and reconciliation: schedule(), reconcile(), and autoscale()
  • Task execution: exec_in_container(), get_process_status(), and profile_task()
  • Lifecycle management: bind_runtime(), seed_liveness(), run_teardown(), teardown(), prune_dead_workers(), and close()
  • Resource inspection: status(), autoscaler_status(), resource_capacity(), and runtime_image()

Each method accepts specific request objects (like ScheduleRequest) and returns result objects (like ScheduleResult) that the Iris controller uses to make orchestration decisions.

Step 1: Create the Backend Implementation

Create a new Python module at lib/iris/src/iris/cluster/backends/<myprovider>/backend.py. This location follows the convention established by existing providers, such as the reference RPC backend at lib/iris/src/iris/cluster/backends/rpc/backend.py.

Define the BackendDescriptor

Every backend must expose a class attribute named descriptor containing a BackendDescriptor instance. This metadata tells Iris how to identify your backend, what kind of cluster it manages, and what resources it advertises to the scheduler.

from iris.cluster.controller.backend import (
    BackendDescriptor,
    BackendKind,
)

class MyProviderBackend:
    descriptor = BackendDescriptor(
        backend_id="myprovider",
        kind=BackendKind.WORKER,  # or BackendKind.KUBERNETES

        advertised_attributes={"accelerator": {"h100"}},
        scale_groups=frozenset({"default"}),
        display_name="MyProvider Cluster",
    )

The kind parameter must be either BackendKind.WORKER for direct worker management or BackendKind.KUBERNETES for Kubernetes-based orchestration. The advertised_attributes dictionary communicates hardware capabilities to Iris's scheduling engine.

Implement the Protocol Methods

Your class must implement all 15 abstract methods from the TaskBackend protocol. The controller first calls bind_runtime() to inject the BackendRuntime instance, then uses the remaining methods to manage workload lifecycles.

from iris.cluster.controller.backend import (
    TaskBackend,
    BackendRuntime,
    ScheduleRequest,
    ScheduleResult,
    ReconcileRequest,
    ReconcileResult,
    AutoscaleRequest,
    AutoscaleResult,
)

class MyProviderBackend:
    def __init__(self) -> None:
        self._runtime: BackendRuntime | None = None
        self.autoscaler = None  # Set to Autoscaler instance if managing capacity

    def bind_runtime(self, runtime: BackendRuntime) -> None:
        self._runtime = runtime

    def schedule(self, request: ScheduleRequest) -> ScheduleResult:
        # Pure decision logic - delegate to iris.cluster.controller.scheduling

        return ScheduleResult()

    def reconcile(self, request: ReconcileRequest) -> ReconcileResult:
        # Communicate with underlying cluster, translate to ControllerEffects

        return ReconcileResult()

    def autoscale(self, request: AutoscaleRequest) -> AutoscaleResult:
        # Provision or tear down resources based on demand

        return AutoscaleResult()

    def status(self):
        # Return controller_pb2.Controller.BackendStatus for the dashboard

        pass

Implement the remaining methods (get_process_status, exec_in_container, resource_capacity, runtime_image, autoscaler_status, seed_liveness, run_teardown, teardown, prune_dead_workers, and close) following the same pattern of accepting request objects and returning structured results.

Step 2: Register the Provider in the Factory

Wire your backend into Iris's discovery mechanism by updating the registry in lib/iris/src/iris/cluster/backends/__init__.py. Add an entry mapping your provider identifier to the backend class.

from .rpc.backend import RpcBackend
from .myprovider.backend import MyProviderBackend

BACKEND_REGISTRY = {
    "rpc": RpcBackend,
    "myprovider": MyProviderBackend,
}

The controller consults this dictionary when the --backend configuration specifies your provider identifier. The mapping enables Iris to instantiate your class and begin calling protocol methods.

Step 3: Update CLI and Configuration

Expose the new provider through Iris's command-line interface by modifying lib/iris/src/iris/cli.py. Add your identifier to the allowed choices for the --backend argument.

parser.add_argument(
    "--backend",
    choices=list(BACKEND_REGISTRY.keys()),
    default="rpc",
    help="Cluster backend provider to use.",
)

If your backend requires custom configuration (such as image names, API endpoints, or resource limits), expose these as additional CLI flags or as fields in your BackendDescriptor. Parse these values in your backend's __init__ method or bind_runtime implementation.

Step 4: Implement Platform-Specific Logic

For backends communicating via RPC or proprietary APIs, study the reference implementation in lib/iris/src/iris/cluster/backends/rpc/backend.py. This file demonstrates concrete patterns for:

  • Handling RPC request/response structures
  • Managing connection lifecycle and error handling
  • Translating cluster-specific states into Iris ControllerEffects
  • Implementing graceful teardown and worker pruning

Copy the structural patterns from this file while replacing the RPC-specific logic with your platform's SDK or API calls.

Step 5: Add Comprehensive Tests

Validate your implementation with unit and integration tests. Mirror the style of existing tests under lib/iris/tests/, particularly the end-to-end backend tests at lib/iris/tests/e2e/test_backend.py.

Write unit tests that exercise each protocol method in isolation, mocking the underlying cluster API. Add integration tests that spin up a minimal instance of your backend (or a mock server) and verify that Iris can successfully schedule tasks, reconcile state, and trigger autoscale operations.

Step 6: Document the Provider

Update the Iris README or create documentation in the docs/ folder describing your backend. Include required configuration parameters, authentication setup, advertised attributes, and any platform-specific limitations or behaviors.

Summary

Frequently Asked Questions

What methods are required to implement the TaskBackend protocol?

You must implement 15 abstract methods defined in lib/iris/src/iris/cluster/controller/backend.py: schedule, reconcile, autoscale, status, autoscaler_status, resource_capacity, runtime_image, get_process_status, profile_task, exec_in_container, bind_runtime, seed_liveness, run_teardown, teardown, prune_dead_workers, and close. These methods handle the complete lifecycle from task scheduling through resource cleanup.

How do I register a new cluster backend provider with Iris?

Import your backend class in lib/iris/src/iris/cluster/backends/__init__.py and add it to the BACKEND_REGISTRY dictionary, mapping a string identifier (like "myprovider") to your class. The Iris controller uses this mapping to instantiate your backend when the user specifies the --backend flag or configuration option.

What is the difference between BackendKind.WORKER and BackendKind.KUBERNETES?

BackendKind.WORKER indicates that your backend manages individual worker processes or VMs directly, while BackendKind.KUBERNETES signals that your backend integrates with a Kubernetes cluster. This distinction helps Iris apply the correct scheduling heuristics and lifecycle management policies for your infrastructure type.

Where can I find a reference implementation for guidance?

The RPC backend at lib/iris/src/iris/cluster/backends/rpc/backend.py provides a complete, production-ready implementation. Study this file for concrete examples of request handling, error management, and the proper implementation of each protocol method including reconcile and autoscale.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →