# What Is Fray ResourceConfig and How Does It Enable Distributed Execution?

> Learn how Fray ResourceConfig enables distributed execution by declaratively defining hardware requirements for tasks in marin-community/marin jobs.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: internals
- Published: 2026-08-29

---

**Fray `ResourceConfig` is the central data structure that declaratively specifies hardware requirements—including CPU, RAM, accelerators, and gang-scheduling constraints—for tasks in distributed jobs within the marin-community/marin repository.**

In the marin-community/marin ecosystem, **Fray ResourceConfig** serves as the foundational abstraction for defining resource requirements in distributed machine learning workloads. This Python class encapsulates everything from basic CPU and memory allocations to complex multi-device TPU slices, enabling developers to express infrastructure needs without managing low-level orchestration details. When a `JobRequest` is submitted, the attached `ResourceConfig` determines exactly how and where the task executes across the cluster.

## Core Purpose of ResourceConfig in Fray

The `ResourceConfig` class, defined in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py) (lines 12-30), acts as the single source of truth for task resource requirements. It captures four critical dimensions of execution context:

- **Compute Resources**: CPU cores, RAM allocation, and disk size
- **Accelerator Hardware**: Device type (CPU, GPU, or TPU) with specific model specifications
- **Scheduling Semantics**: Gang-scheduling information including replica counts for distributed training
- **Placement Policies**: Region constraints, pre-emptibility settings, and target cluster routing

When you create a `JobRequest`, the `resources` parameter accepts a `ResourceConfig` instance that travels through [`lib/fray/src/fray/client.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/client.py) to the backend implementation in [`lib/fray/src/fray/iris_backend.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/iris_backend.py). The backend uses this configuration to allocate appropriate hardware, set environment variables, and enforce placement constraints.

## Hardware Specification and Accelerator Support

Fray provides specialized factory methods for common accelerator types while supporting direct instantiation for custom configurations.

### TPU and GPU Configuration

For TPU workloads, `ResourceConfig.with_tpu()` constructs configurations that specify slice sizes like `v4-8`. The backend uses these specifications to request multi-host TPU slices and populate accelerator-specific environment variables such as `JAX_PLATFORMS` and `LIBTPU_INIT_ARGS`.

```python

# Request a single-slice TPU (v4-8) for training

rc = ResourceConfig.with_tpu("v4-8")
job = JobRequest(
    name="train-tpu",
    entrypoint=Entrypoint.from_callable(train),
    resources=rc
)

```

For GPU clusters, `ResourceConfig.with_gpu()` accepts device types like `H100` and replica counts. As demonstrated in [`tests/inference/test_serve.py`](https://github.com/marin-community/marin/blob/main/tests/inference/test_serve.py) (lines 173-176), this pattern supports multi-GPU inference deployments.

```python

# Request 8 H100 GPUs with custom disk allocation

rc = ResourceConfig.with_gpu("H100", count=8, disk="100g")
job = JobRequest(
    name="train-gpu",
    entrypoint=Entrypoint.from_callable(train),
    resources=rc
)

```

### CPU-Only Workloads

When accelerators are unnecessary, direct instantiation provides granular control over compute resources. The test suite in [`lib/zephyr/tests/test_vortex.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/tests/test_vortex.py) (lines 30-33) demonstrates lightweight CPU configurations.

```python

# Simple CPU job with custom container image

rc = ResourceConfig(
    cpu=4,
    ram="16g",
    image="gcr.io/my-project/custom-runner:latest",
    regions=[ANY_REGION],
)
job = JobRequest(
    name="cpu-task",
    entrypoint=Entrypoint.from_callable(task),
    resources=rc
)

```

## Gang Scheduling and Multi-Replica Execution

The `replicas` field within `ResourceConfig` enables **gang scheduling**, a critical feature for distributed training workloads. When you specify multiple replicas—such as for multi-slice TPU training—the scheduler guarantees that all instances start simultaneously or not at all. This prevents resource deadlocks where partial allocations would waste expensive accelerator time.

According to the test patterns in [`tests/test_training.py`](https://github.com/marin-community/marin/blob/main/tests/test_training.py) (lines 11-20), TPU configurations automatically infer replica counts from the slice specification, ensuring that multi-host training jobs maintain synchronization across all nodes.

## Placement and Routing Controls

Beyond hardware allocation, `ResourceConfig` determines geographical and logical placement within the compute fabric.

### Region Constraints and ANY_REGION

The `regions` parameter accepts specific zone identifiers or the special `ANY_REGION` constant. When `ANY_REGION` appears in the list, the task becomes eligible for execution in any available region, overriding parent job constraints. This flexibility proves essential for batch workloads insensitive to data locality, while strict region lists enforce compliance or latency requirements.

### Container Overrides and Pre-emptibility

You can override the default container image per-task using the `image` parameter, enabling heterogeneous workflows where different stages require different dependencies. The configuration also exposes pre-emptibility settings that allow the scheduler to reclaim resources for higher-priority jobs, reducing costs for fault-tolerant training runs.

## Submitting Jobs with ResourceConfig

The submission flow coordinates between client-side configuration and backend orchestration:

1. **Client Validation**: [`lib/fray/src/fray/client.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/client.py) validates the `ResourceConfig` against the `JobRequest` schema
2. **Backend Translation**: [`lib/fray/src/fray/iris_backend.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/iris_backend.py) converts the configuration into cluster-specific resource requests
3. **Environment Injection**: The backend populates runtime variables based on the selected accelerator type
4. **Resource Allocation**: The scheduler reserves the specified hardware, respecting gang-scheduling constraints

```python

# Complete workflow from configuration to submission

from fray.types import ResourceConfig, ANY_REGION
from fray.client import submit

rc = ResourceConfig.with_tpu("v4-8", preemptible=True)
job = JobRequest(
    name="distributed-training",
    entrypoint=Entrypoint.from_callable(train_fn),
    resources=rc
)

# The submit method forwards ResourceConfig to Iris backend

submit(job)

```

## Summary

- **Fray ResourceConfig**, defined in [`lib/fray/src/fray/types.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/types.py), declaratively captures CPU, RAM, disk, accelerator, and scheduling requirements for distributed tasks.
- Factory methods `with_tpu()` and `with_gpu()` simplify hardware specification while direct instantiation supports custom CPU and container configurations.
- The `replicas` field enables gang scheduling, ensuring multi-slice TPU or multi-GPU jobs start atomically across the cluster.
- Placement controls including `regions`, `ANY_REGION`, and pre-emptibility policies allow fine-grained control over execution location and cost.
- The configuration flows from `JobRequest` through [`lib/fray/src/fray/client.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/client.py) to [`lib/fray/src/fray/iris_backend.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/iris_backend.py), where it materializes into allocated resources and environment variables.

## Frequently Asked Questions

### How does Fray ResourceConfig handle TPU slice requests?

Fray uses `ResourceConfig.with_tpu()` to parse slice specifications like `v4-8` and automatically configures the appropriate replica count for gang scheduling. According to [`tests/test_training.py`](https://github.com/marin-community/marin/blob/main/tests/test_training.py), the backend translates these configurations into environment variables such as `JAX_PLATFORMS` and `LIBTPU_INIT_ARGS` required for JAX-based training.

### What is the difference between `regions` and `ANY_REGION` in Fray ResourceConfig?

Specific region strings in the `regions` list constrain execution to particular geographical zones, while `ANY_REGION` serves as a wildcard that allows the scheduler to place the task in any available region. Using `ANY_REGION` effectively opts out of parent job region constraints for maximum scheduling flexibility.

### How does `ResourceConfig` enable gang scheduling for distributed training?

The `replicas` field specifies how many identical task instances must launch simultaneously. When submitting multi-slice TPU jobs, this ensures all slices start together or not at all, preventing partial allocations that would stall distributed training. The Iris backend uses this field to reserve the complete set of resources before starting execution.

### Can I override container images per task using `ResourceConfig`?

Yes. The `image` parameter in `ResourceConfig` allows you to specify custom container registries such as `gcr.io/my-project/custom-runner:latest`. This override applies only to the specific task, enabling workflows where preprocessing, training, and inference stages each require distinct dependencies without modifying the global job configuration.