# How to Add a New Cluster Backend Provider to Iris: Complete Implementation Guide

> Learn how to add a new cluster backend provider to Iris. This guide covers implementing TaskBackend, registering your provider, and updating the CLI for seamless integration. Enhance Iris functionality today.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**To add a new cluster backend provider to Iris, implement the `TaskBackend` protocol in a new module under `lib/iris/src/iris/cluster/backends/`, register the class in the `BACKEND_REGISTRY` dictionary in [`lib/iris/src/iris/cluster/backends/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/__init__.py), and update the CLI parser in [`lib/iris/src/iris/cli.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cli.py) to recognize the new provider identifier.**

Iris, the distributed task orchestrator in the marin-community/marin repository, uses a pluggable backend architecture to abstract cluster operations. When you add a new cluster backend provider to Iris, you enable the controller to schedule, reconcile, and autoscale workloads on your custom infrastructure. This guide walks through the protocol implementation, registration process, and integration steps required to extend Iris with a new backend.

## Understand the TaskBackend Protocol

The `TaskBackend` protocol defines the contract between Iris and cluster infrastructure. Defined in [`lib/iris/src/iris/cluster/controller/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/controller/backend.py), it requires 15 concrete implementations that handle the complete lifecycle of task scheduling, execution, and resource management.

Your implementation must provide these core operational methods:

- **Scheduling and reconciliation**: `schedule()`, `reconcile()`, and `autoscale()`
- **Task execution**: `exec_in_container()`, `get_process_status()`, and `profile_task()`
- **Lifecycle management**: `bind_runtime()`, `seed_liveness()`, `run_teardown()`, `teardown()`, `prune_dead_workers()`, and `close()`
- **Resource inspection**: `status()`, `autoscaler_status()`, `resource_capacity()`, and `runtime_image()`

Each method accepts specific request objects (like `ScheduleRequest`) and returns result objects (like `ScheduleResult`) that the Iris controller uses to make orchestration decisions.

## Step 1: Create the Backend Implementation

Create a new Python module at `lib/iris/src/iris/cluster/backends/<myprovider>/backend.py`. This location follows the convention established by existing providers, such as the reference RPC backend at [`lib/iris/src/iris/cluster/backends/rpc/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/rpc/backend.py).

### Define the BackendDescriptor

Every backend must expose a class attribute named `descriptor` containing a `BackendDescriptor` instance. This metadata tells Iris how to identify your backend, what kind of cluster it manages, and what resources it advertises to the scheduler.

```python
from iris.cluster.controller.backend import (
    BackendDescriptor,
    BackendKind,
)

class MyProviderBackend:
    descriptor = BackendDescriptor(
        backend_id="myprovider",
        kind=BackendKind.WORKER,  # or BackendKind.KUBERNETES

        advertised_attributes={"accelerator": {"h100"}},
        scale_groups=frozenset({"default"}),
        display_name="MyProvider Cluster",
    )

```

The `kind` parameter must be either `BackendKind.WORKER` for direct worker management or `BackendKind.KUBERNETES` for Kubernetes-based orchestration. The `advertised_attributes` dictionary communicates hardware capabilities to Iris's scheduling engine.

### Implement the Protocol Methods

Your class must implement all 15 abstract methods from the `TaskBackend` protocol. The controller first calls `bind_runtime()` to inject the `BackendRuntime` instance, then uses the remaining methods to manage workload lifecycles.

```python
from iris.cluster.controller.backend import (
    TaskBackend,
    BackendRuntime,
    ScheduleRequest,
    ScheduleResult,
    ReconcileRequest,
    ReconcileResult,
    AutoscaleRequest,
    AutoscaleResult,
)

class MyProviderBackend:
    def __init__(self) -> None:
        self._runtime: BackendRuntime | None = None
        self.autoscaler = None  # Set to Autoscaler instance if managing capacity

    def bind_runtime(self, runtime: BackendRuntime) -> None:
        self._runtime = runtime

    def schedule(self, request: ScheduleRequest) -> ScheduleResult:
        # Pure decision logic - delegate to iris.cluster.controller.scheduling

        return ScheduleResult()

    def reconcile(self, request: ReconcileRequest) -> ReconcileResult:
        # Communicate with underlying cluster, translate to ControllerEffects

        return ReconcileResult()

    def autoscale(self, request: AutoscaleRequest) -> AutoscaleResult:
        # Provision or tear down resources based on demand

        return AutoscaleResult()

    def status(self):
        # Return controller_pb2.Controller.BackendStatus for the dashboard

        pass

```

Implement the remaining methods (`get_process_status`, `exec_in_container`, `resource_capacity`, `runtime_image`, `autoscaler_status`, `seed_liveness`, `run_teardown`, `teardown`, `prune_dead_workers`, and `close`) following the same pattern of accepting request objects and returning structured results.

## Step 2: Register the Provider in the Factory

Wire your backend into Iris's discovery mechanism by updating the registry in [`lib/iris/src/iris/cluster/backends/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/__init__.py). Add an entry mapping your provider identifier to the backend class.

```python
from .rpc.backend import RpcBackend
from .myprovider.backend import MyProviderBackend

BACKEND_REGISTRY = {
    "rpc": RpcBackend,
    "myprovider": MyProviderBackend,
}

```

The controller consults this dictionary when the `--backend` configuration specifies your provider identifier. The mapping enables Iris to instantiate your class and begin calling protocol methods.

## Step 3: Update CLI and Configuration

Expose the new provider through Iris's command-line interface by modifying [`lib/iris/src/iris/cli.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cli.py). Add your identifier to the allowed choices for the `--backend` argument.

```python
parser.add_argument(
    "--backend",
    choices=list(BACKEND_REGISTRY.keys()),
    default="rpc",
    help="Cluster backend provider to use.",
)

```

If your backend requires custom configuration (such as image names, API endpoints, or resource limits), expose these as additional CLI flags or as fields in your `BackendDescriptor`. Parse these values in your backend's `__init__` method or `bind_runtime` implementation.

## Step 4: Implement Platform-Specific Logic

For backends communicating via RPC or proprietary APIs, study the reference implementation in [`lib/iris/src/iris/cluster/backends/rpc/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/rpc/backend.py). This file demonstrates concrete patterns for:

- Handling RPC request/response structures
- Managing connection lifecycle and error handling
- Translating cluster-specific states into Iris `ControllerEffects`
- Implementing graceful teardown and worker pruning

Copy the structural patterns from this file while replacing the RPC-specific logic with your platform's SDK or API calls.

## Step 5: Add Comprehensive Tests

Validate your implementation with unit and integration tests. Mirror the style of existing tests under `lib/iris/tests/`, particularly the end-to-end backend tests at [`lib/iris/tests/e2e/test_backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/tests/e2e/test_backend.py).

Write unit tests that exercise each protocol method in isolation, mocking the underlying cluster API. Add integration tests that spin up a minimal instance of your backend (or a mock server) and verify that Iris can successfully schedule tasks, reconcile state, and trigger autoscale operations.

## Step 6: Document the Provider

Update the Iris README or create documentation in the `docs/` folder describing your backend. Include required configuration parameters, authentication setup, advertised attributes, and any platform-specific limitations or behaviors.

## Summary

- Implement all 15 methods of the `TaskBackend` protocol defined in [`lib/iris/src/iris/cluster/controller/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/controller/backend.py).
- Create a `BackendDescriptor` with unique `backend_id`, appropriate `kind`, and `advertised_attributes` for scheduling.
- Register your class in `BACKEND_REGISTRY` inside [`lib/iris/src/iris/cluster/backends/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/__init__.py).
- Update [`lib/iris/src/iris/cli.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cli.py) to include the new provider in `--backend` choices.
- Reference [`lib/iris/src/iris/cluster/backends/rpc/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/rpc/backend.py) for implementation patterns and error handling.
- Add unit and integration tests under `lib/iris/tests/` to verify scheduling, reconciliation, and autoscaling behavior.

## Frequently Asked Questions

### What methods are required to implement the TaskBackend protocol?

You must implement 15 abstract methods defined in [`lib/iris/src/iris/cluster/controller/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/controller/backend.py): `schedule`, `reconcile`, `autoscale`, `status`, `autoscaler_status`, `resource_capacity`, `runtime_image`, `get_process_status`, `profile_task`, `exec_in_container`, `bind_runtime`, `seed_liveness`, `run_teardown`, `teardown`, `prune_dead_workers`, and `close`. These methods handle the complete lifecycle from task scheduling through resource cleanup.

### How do I register a new cluster backend provider with Iris?

Import your backend class in [`lib/iris/src/iris/cluster/backends/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/__init__.py) and add it to the `BACKEND_REGISTRY` dictionary, mapping a string identifier (like `"myprovider"`) to your class. The Iris controller uses this mapping to instantiate your backend when the user specifies the `--backend` flag or configuration option.

### What is the difference between BackendKind.WORKER and BackendKind.KUBERNETES?

`BackendKind.WORKER` indicates that your backend manages individual worker processes or VMs directly, while `BackendKind.KUBERNETES` signals that your backend integrates with a Kubernetes cluster. This distinction helps Iris apply the correct scheduling heuristics and lifecycle management policies for your infrastructure type.

### Where can I find a reference implementation for guidance?

The RPC backend at [`lib/iris/src/iris/cluster/backends/rpc/backend.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/cluster/backends/rpc/backend.py) provides a complete, production-ready implementation. Study this file for concrete examples of request handling, error management, and the proper implementation of each protocol method including `reconcile` and `autoscale`.