What Is Fray ResourceConfig and How Does It Enable Distributed Execution?
Fray ResourceConfig is the central data structure that declaratively specifies hardware requirements—including CPU, RAM, accelerators, and gang-scheduling constraints—for tasks in distributed jobs within the marin-community/marin repository.
In the marin-community/marin ecosystem, Fray ResourceConfig serves as the foundational abstraction for defining resource requirements in distributed machine learning workloads. This Python class encapsulates everything from basic CPU and memory allocations to complex multi-device TPU slices, enabling developers to express infrastructure needs without managing low-level orchestration details. When a JobRequest is submitted, the attached ResourceConfig determines exactly how and where the task executes across the cluster.
Core Purpose of ResourceConfig in Fray
The ResourceConfig class, defined in lib/fray/src/fray/types.py (lines 12-30), acts as the single source of truth for task resource requirements. It captures four critical dimensions of execution context:
- Compute Resources: CPU cores, RAM allocation, and disk size
- Accelerator Hardware: Device type (CPU, GPU, or TPU) with specific model specifications
- Scheduling Semantics: Gang-scheduling information including replica counts for distributed training
- Placement Policies: Region constraints, pre-emptibility settings, and target cluster routing
When you create a JobRequest, the resources parameter accepts a ResourceConfig instance that travels through lib/fray/src/fray/client.py to the backend implementation in lib/fray/src/fray/iris_backend.py. The backend uses this configuration to allocate appropriate hardware, set environment variables, and enforce placement constraints.
Hardware Specification and Accelerator Support
Fray provides specialized factory methods for common accelerator types while supporting direct instantiation for custom configurations.
TPU and GPU Configuration
For TPU workloads, ResourceConfig.with_tpu() constructs configurations that specify slice sizes like v4-8. The backend uses these specifications to request multi-host TPU slices and populate accelerator-specific environment variables such as JAX_PLATFORMS and LIBTPU_INIT_ARGS.
# Request a single-slice TPU (v4-8) for training
rc = ResourceConfig.with_tpu("v4-8")
job = JobRequest(
name="train-tpu",
entrypoint=Entrypoint.from_callable(train),
resources=rc
)
For GPU clusters, ResourceConfig.with_gpu() accepts device types like H100 and replica counts. As demonstrated in tests/inference/test_serve.py (lines 173-176), this pattern supports multi-GPU inference deployments.
# Request 8 H100 GPUs with custom disk allocation
rc = ResourceConfig.with_gpu("H100", count=8, disk="100g")
job = JobRequest(
name="train-gpu",
entrypoint=Entrypoint.from_callable(train),
resources=rc
)
CPU-Only Workloads
When accelerators are unnecessary, direct instantiation provides granular control over compute resources. The test suite in lib/zephyr/tests/test_vortex.py (lines 30-33) demonstrates lightweight CPU configurations.
# Simple CPU job with custom container image
rc = ResourceConfig(
cpu=4,
ram="16g",
image="gcr.io/my-project/custom-runner:latest",
regions=[ANY_REGION],
)
job = JobRequest(
name="cpu-task",
entrypoint=Entrypoint.from_callable(task),
resources=rc
)
Gang Scheduling and Multi-Replica Execution
The replicas field within ResourceConfig enables gang scheduling, a critical feature for distributed training workloads. When you specify multiple replicas—such as for multi-slice TPU training—the scheduler guarantees that all instances start simultaneously or not at all. This prevents resource deadlocks where partial allocations would waste expensive accelerator time.
According to the test patterns in tests/test_training.py (lines 11-20), TPU configurations automatically infer replica counts from the slice specification, ensuring that multi-host training jobs maintain synchronization across all nodes.
Placement and Routing Controls
Beyond hardware allocation, ResourceConfig determines geographical and logical placement within the compute fabric.
Region Constraints and ANY_REGION
The regions parameter accepts specific zone identifiers or the special ANY_REGION constant. When ANY_REGION appears in the list, the task becomes eligible for execution in any available region, overriding parent job constraints. This flexibility proves essential for batch workloads insensitive to data locality, while strict region lists enforce compliance or latency requirements.
Container Overrides and Pre-emptibility
You can override the default container image per-task using the image parameter, enabling heterogeneous workflows where different stages require different dependencies. The configuration also exposes pre-emptibility settings that allow the scheduler to reclaim resources for higher-priority jobs, reducing costs for fault-tolerant training runs.
Submitting Jobs with ResourceConfig
The submission flow coordinates between client-side configuration and backend orchestration:
- Client Validation:
lib/fray/src/fray/client.pyvalidates theResourceConfigagainst theJobRequestschema - Backend Translation:
lib/fray/src/fray/iris_backend.pyconverts the configuration into cluster-specific resource requests - Environment Injection: The backend populates runtime variables based on the selected accelerator type
- Resource Allocation: The scheduler reserves the specified hardware, respecting gang-scheduling constraints
# Complete workflow from configuration to submission
from fray.types import ResourceConfig, ANY_REGION
from fray.client import submit
rc = ResourceConfig.with_tpu("v4-8", preemptible=True)
job = JobRequest(
name="distributed-training",
entrypoint=Entrypoint.from_callable(train_fn),
resources=rc
)
# The submit method forwards ResourceConfig to Iris backend
submit(job)
Summary
- Fray ResourceConfig, defined in
lib/fray/src/fray/types.py, declaratively captures CPU, RAM, disk, accelerator, and scheduling requirements for distributed tasks. - Factory methods
with_tpu()andwith_gpu()simplify hardware specification while direct instantiation supports custom CPU and container configurations. - The
replicasfield enables gang scheduling, ensuring multi-slice TPU or multi-GPU jobs start atomically across the cluster. - Placement controls including
regions,ANY_REGION, and pre-emptibility policies allow fine-grained control over execution location and cost. - The configuration flows from
JobRequestthroughlib/fray/src/fray/client.pytolib/fray/src/fray/iris_backend.py, where it materializes into allocated resources and environment variables.
Frequently Asked Questions
How does Fray ResourceConfig handle TPU slice requests?
Fray uses ResourceConfig.with_tpu() to parse slice specifications like v4-8 and automatically configures the appropriate replica count for gang scheduling. According to tests/test_training.py, the backend translates these configurations into environment variables such as JAX_PLATFORMS and LIBTPU_INIT_ARGS required for JAX-based training.
What is the difference between regions and ANY_REGION in Fray ResourceConfig?
Specific region strings in the regions list constrain execution to particular geographical zones, while ANY_REGION serves as a wildcard that allows the scheduler to place the task in any available region. Using ANY_REGION effectively opts out of parent job region constraints for maximum scheduling flexibility.
How does ResourceConfig enable gang scheduling for distributed training?
The replicas field specifies how many identical task instances must launch simultaneously. When submitting multi-slice TPU jobs, this ensures all slices start together or not at all, preventing partial allocations that would stall distributed training. The Iris backend uses this field to reserve the complete set of resources before starting execution.
Can I override container images per task using ResourceConfig?
Yes. The image parameter in ResourceConfig allows you to specify custom container registries such as gcr.io/my-project/custom-runner:latest. This override applies only to the specific task, enabling workflows where preprocessing, training, and inference stages each require distinct dependencies without modifying the global job configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →