# How Fray's ResourceConfig Allocates Distributed Compute Resources for Marin

> Learn how Fray's ResourceConfig efficiently allocates distributed compute resources for Marin by converting task requirements into Iris scheduler protobufs for cluster replica distribution.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-28

---

**Fray's `ResourceConfig` dataclass captures per-task compute requirements—CPU, RAM, disk, and accelerator type—then converts these into Iris scheduler protobufs to distribute replicas across Marin's compute clusters.**

The Marin distributed execution framework uses Fray to translate high-level job definitions into cluster-wide resource allocations. At the core of this system lies the immutable **`ResourceConfig`** dataclass defined in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py), which declaratively expresses the compute needs of individual task replicas before the Iris backend translates them into distributed placements.

## ResourceConfig Dataclass and Per-Task Specification

The `ResourceConfig` dataclass in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py) serves as the canonical representation of resources required for a single replica of a Marin job. It encapsulates:

- **Compute cores**: CPU count and RAM allocation
- **Storage**: Temporary disk requirements
- **Accelerator type**: Device specification (CPU, GPU, or TPU)
- **Scheduling hints**: Pre-emptibility flags, region/zone constraints, custom container images, and alternative device variants

This immutable structure ensures that resource specifications remain consistent throughout the job lifecycle, from submission to execution.

## Device-Specific Configuration Builders

Fray provides three static factory methods on `ResourceConfig` to simplify construction for common accelerator types.

### CPU Configuration

The `with_cpu()` method creates a pure-CPU configuration without accelerator devices. This is the baseline for jobs requiring only general-purpose compute.

```python
from fray import ResourceConfig, JobRequest, Entrypoint, create_environment

resources = ResourceConfig.with_cpu(cpu=4, ram="16g")
env = create_environment(workspace="my_proj")
job = JobRequest(
    name="cpu_job",
    entrypoint=Entrypoint.from_callable(my_train_fn),
    resources=resources,
    environment=env,
)
client.submit(job)

```

### GPU Configuration

The `with_gpu(gpu_type, count)` method constructs configurations for NVIDIA GPUs. It stores device specifications in a **`GpuConfig`** nested dataclass (defined at lines 40-48 of [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py)), capturing both the GPU model and quantity required.

```python
resources = ResourceConfig.with_gpu("A100-80G", count=8)
job = JobRequest(
    name="gpu_job",
    entrypoint=Entrypoint.from_binary("python", ["train.py"]),
    resources=resources,
)
client.submit(job)

```

### TPU Configuration

The `with_tpu(tpu_type, slice_count=...)` method handles TPU slice topology. Implemented at lines 96-104 of [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py), this builder:

1. Calculates the total number of replicas from the chip topology and requested slice count
2. Defaults host CPU and RAM to **50%** of the TPU VM's host resources using the `TPU_HOST_RESOURCES` constant and `DEFAULT_TPU_HOST_FRACTION` parameter

```python
resources = ResourceConfig.with_tpu("v5p-8", slice_count=2)
job = JobRequest(
    name="tpu_job",
    entrypoint=Entrypoint.from_binary("python", ["train.py"]),
    resources=resources,
)
client.submit(job)

```

## Translation to Iris Scheduler Protobufs

When a job is submitted, Fray's backend translates `ResourceConfig` instances into Iris scheduler structures via [`lib/fray/src/fray/iris_backend.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/iris_backend.py).

### Resource Specifications

The `convert_resources()` function (lines 17-33) maps `ResourceConfig` fields to Iris `ResourceSpec` protobufs:

- `cpu` and `ram` map directly to compute allocations
- `disk` translates to temporary storage requirements
- Device configurations (`GpuConfig` or TPU specs) convert to accelerator descriptions

### Scheduling Constraints

The `convert_constraints()` function (lines 35-64) derives Iris scheduling constraints from `ResourceConfig` metadata:

- Pre-emptibility flags
- Region and zone restrictions
- Target cluster specifications
- Device variant alternatives

## Distributed Allocation and Coscheduling

Iris performs the actual placement of replicas across Marin's clusters using the translated specifications. The allocation process receives:

- **Per-replica `ResourceSpec`** containing CPU, RAM, disk, and device requirements
- **Replica count** derived from `resources.replicas`, where TPU jobs calculate total replicas as `slice_count * vm_count`
- **Coscheduling groups** generated by `resolve_coscheduling(resources, replicas)` (called around line 84 of [`iris_backend.py`](https://github.com/marin-community/marin/blob/main/iris_backend.py)), which groups replicas that must share a physical host for multi-GPU or multi-TPU configurations

This architecture separates the *what* (declarative resource requirements) from the *where* (physical cluster placement), allowing Iris to optimize distribution across available hardware while honoring locality constraints.

## Runtime Environment Configuration

For jobs requiring accelerators, Fray merges device-specific environment variables into the job's `EnvironmentSpec` via `convert_environment()` (lines 22-30 of [`iris_backend.py`](https://github.com/marin-community/marin/blob/main/iris_backend.py)). These include:

- **`JAX_PLATFORMS`** for GPU targets
- **`LIBTPU_INIT_ARGS`** for TPU configurations

This ensures that the runtime context matches the allocated hardware without manual configuration by the user.

## Summary

- **`ResourceConfig`** in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py) declaratively captures per-replica compute, memory, disk, and accelerator requirements for Marin jobs.
- **Builder methods** `with_cpu()`, `with_gpu()`, and `with_tpu()` simplify configuration for specific device types, with TPU logic defaulting host resources to 50% of VM capacity via `DEFAULT_TPU_HOST_FRACTION`.
- **`convert_resources()`** and **`convert_constraints()`** in [`iris_backend.py`](https://github.com/marin-community/marin/blob/main/iris_backend.py) translate dataclass fields into Iris scheduler protobufs for cluster-wide distribution.
- **`resolve_coscheduling()`** ensures multi-device replicas are grouped on shared hosts when necessary for performance.
- **Environment variables** like `JAX_PLATFORMS` are automatically injected based on the selected device type to configure the runtime correctly.

## Frequently Asked Questions

### How does ResourceConfig determine the number of TPU replicas?

The `with_tpu()` method calculates the total replica count by multiplying the requested `slice_count` by the number of VMs required for the specified TPU topology. For example, requesting 2 slices of `v5p-8` results in `slice_count * vm_count` total replicas, with host CPU and RAM defaults set via `DEFAULT_TPU_HOST_FRACTION` (50%) in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py) lines 96-104.

### What happens during the conversion from ResourceConfig to Iris specifications?

The `convert_resources()` function (lines 17-33 of [`lib/fray/src/fray/iris_backend.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/iris_backend.py)) maps CPU, RAM, and disk fields to an Iris `ResourceSpec`, while `convert_constraints()` (lines 35-64) translates pre-emptibility, region, zone, and device variant constraints into Iris scheduling directives that determine eligible cluster nodes.

### How does Fray handle multi-GPU allocation on single hosts?

Fray calls `resolve_coscheduling(resources, replicas)` around line 84 of [`iris_backend.py`](https://github.com/marin-community/marin/blob/main/iris_backend.py) to identify replicas that must share a physical machine. Iris uses this coscheduling information to allocate multi-GPU or multi-TPU jobs on hosts with sufficient accelerator density, ensuring coordinated access to local devices.

### Which environment variables are automatically configured for accelerators?

The `convert_environment()` function injects `JAX_PLATFORMS` for GPU jobs and `LIBTPU_INIT_ARGS` for TPU jobs (lines 22-30 of [`iris_backend.py`](https://github.com/marin-community/marin/blob/main/iris_backend.py)), ensuring the runtime environment matches the allocated hardware without manual intervention from the job submitter.