# Configuration Options for Marin: A Complete Guide to the ML Pipeline Framework

> Explore Marin's configuration options for ML pipelines. Master inference, training, and data processing with immutable dataclasses for seamless model serving. Get the complete guide.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-27

---

**Marin uses immutable dataclasses organized into inference, training, and data-processing groups to control every aspect of model serving and training pipelines.**

Marin is a lightweight, composable framework for data ingestion, model training, and inference. All runtime behavior is governed by a hierarchy of frozen dataclasses defined in `lib/marin/src/marin/`, which can be instantiated via YAML/JSON or passed directly through the Python API. This guide covers the complete set of configuration options for Marin, from vLLM engine parameters to TPU resource allocation.

## Inference Configuration Options

Marin’s inference stack is configured through a nested set of dataclasses that control model serving, engine selection, and request routing.

### ServedModelConfig

The `ServedModelConfig` class in [`lib/marin/src/marin/inference/config.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/inference/config.py) (lines 41-50) defines how a model is exposed via an OpenAI-compatible endpoint.

Key fields include:
- `weights` (str): Path to model artifacts (must be non-empty)
- `revision` (str | None): Specific model revision to load
- `api_model` (str | None): API identifier for the served model
- `dtype` (str): Data type for weights, defaults to `"bfloat16"`
- `max_model_len` (int | None): Maximum sequence length (must be positive if supplied)
- `tensor_parallel_size` (int | None): Number of GPUs for tensor parallelism

### VllmEngineConfig

The `VllmEngineConfig` class (lines 77-86) selects the vLLM runtime and controls compilation caching.

```python
from marin.inference.config import VllmEngineConfig, VllmLauncherType, VllmSource

engine_cfg = VllmEngineConfig(
    launcher=VllmLauncherType.CUDA,
    source=VllmSource.MARIN_FORK,
    version="0.25.1",
    compilation_cache=VllmCompilationCacheMode.MANAGED,
    startup_timeout_seconds=1800,
    extra_args=("--enforce-eager",)
)

```

Validation enforces that `startup_timeout_seconds` must be greater than 0, and using `source=VllmSource.MARIN_FORK` requires `launcher=VllmLauncherType.CUDA`.

### InferenceProxyConfig and InferenceWorkerConfig

Network behavior is controlled by `InferenceProxyConfig` (lines 117-124) and `InferenceWorkerConfig` (lines 142-149):

- `InferenceProxyConfig` manages the HTTP front-end with fields like `port`, `request_timeout_seconds` (default 300.0), `max_pending_requests` (default 256), and `response_fetch_batch_size` (default 64)
- `InferenceWorkerConfig` limits per-worker concurrency via `max_in_flight` (default 16) and `request_timeout_seconds` (default 180.0)

All timeout and count values must be positive, and ports cannot be negative.

### BrokerConfig and Iris Placement

For distributed inference, `BrokerConfig` (lines 154-169) multiplexes workers and proxies. It bundles worker and proxy configurations while enforcing that the broker must run on a non-preemptible resource. The config validates that timeout values satisfy the hierarchy: `0 < worker < lease < proxy`.

`IrisConfig` (lines 92-102) handles remote inference placement through the Iris job-orchestration layer, specifying `worker_resources`, `cache_ttl_days` (default 14), and retry policies. TPU instances configured here must use single-host VMs.

### RemoteInferenceConfig

The top-level `RemoteInferenceConfig` (lines 227-236) composes model, engine, and placement:

```python
from marin.inference.config import (
    ServedModelConfig, VllmEngineConfig, RemoteInferenceConfig, IrisConfig
)
from fray.types import ResourceConfig, GpuConfig, EnvironmentConfig

remote_cfg = RemoteInferenceConfig(
    model=ServedModelConfig(
        weights="gs://my-bucket/llama-7b",
        api_model="llama-7b",
        dtype="bfloat16"
    ),
    engine=VllmEngineConfig(launcher=VllmLauncherType.CUDA),
    iris=IrisConfig(
        worker_resources=ResourceConfig.with_gpu(gpu=GpuConfig(num_gpus=1)),
        worker_environment=EnvironmentConfig(),
        cache_ttl_days=7
    ),
    instances=2
)

```

The `instances` field must be positive and determines how many workers Iris spawns.

## Training Configuration Options

Marin wraps the Levanter training library with configuration classes that handle resource allocation and environment setup.

### TrainLmOnPodConfig and TrainDpoOnPodConfig

Defined in [`lib/marin/src/marin/training/training.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/training/training.py) (lines 90-107 and 109-126), these classes manage LM and DPO training on GKE/TPU pods:

```python
from marin.training.training import TrainLmOnPodConfig
from levanter.main.train_lm import TrainLmConfig
from fray.types import ResourceConfig, TpuConfig

pod_cfg = TrainLmOnPodConfig(
    train_config=TrainLmConfig(model=model_def, trainer=trainer_cfg),
    resources=ResourceConfig.with_tpu(
        tpu=TpuConfig(variant="v4-8", vm_count=1)
    ),
    output_path="gs://my-bucket/experiments/run-123",
    env_vars={"WANDB_API_KEY": "..."},
    auto_build_caches=False  # Default: prevents accidental cache creation

)

```

`TrainDpoOnPodConfig` adds `auto_num_epochs` and `auto_validation_runs` fields that auto-compute training steps when specified.

### Resource Allocation

The `ResourceConfig` class from `fray.types` (referenced in training configs) describes CPU, GPU, and TPU allocation with fields for `cpu`, `ram`, `disk`, `gpu`, `tpu`, and `preemptible` flags. Validation ensures TPU configurations include at least one accelerator.

## Data Ingestion and Transformation Configuration

Marin’s Datakit and Transform sub-packages expose specialized configurations for pipeline preprocessing.

### Datakit Configuration

`SampleCapConfig` in [`lib/marin/src/marin/datakit/ingestion_manifest.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/datakit/ingestion_manifest.py) (lines 71-77) controls ingestion limits:

```python
from marin.datakit.ingestion_manifest import SampleCapConfig

cap_cfg = SampleCapConfig(
    max_bytes=1e9,
    max_examples=10000
)

```

### Transform Pipeline Configuration

The transform layer includes format-specific configurations:
- `WebDataCommonsStagingConfig` ([`lib/marin/src/marin/transform/structured_text/web_data_commons.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/transform/structured_text/web_data_commons.py), lines 49-58): Controls staging of WebDataCommons corpus with bucket and sample cap settings
- `ResiliparseConfig` ([`lib/marin/src/marin/schemas/web/convert.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/schemas/web/convert.py), lines 51-55): Configures HTML-to-Markdown conversion behavior
- `StackExchangeExtractionConfig` ([`lib/marin/src/marin/transform/stackexchange/transform_stackexchange.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/transform/stackexchange/transform_stackexchange.py), lines 28-34): Governs extraction from StackExchange dumps
- `HfRawTextSurfaceConfig`: Manages surface-level processing of HuggingFace text datasets
- `ZeekToDolmaConfig`: Converts Zeek network-security logs into the Dolma schema

## How Configuration Objects Work

All Marin configuration options are implemented as frozen dataclasses (`@dataclass(frozen=True)`), ensuring immutability after construction. This design guarantees reproducibility and enables safe serialization across distributed workers.

The validation layer enforces sensible defaults at instantiation time. For example, `BrokerConfig` validates that timeout values follow a strict hierarchy, while `IrisConfig` ensures TPU placements use single-host VMs. Configurations can be instantiated programmatically, as shown in the examples above, or parsed from YAML/JSON files for CLI usage.

## Summary

- **Immutable dataclasses** form the foundation of Marin configuration, located in [`lib/marin/src/marin/inference/config.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/inference/config.py) and [`lib/marin/src/marin/training/training.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/training/training.py)
- **Inference options** include `ServedModelConfig` for model artifacts, `VllmEngineConfig` for runtime selection, and `RemoteInferenceConfig` for orchestration
- **Training configs** like `TrainLmOnPodConfig` wrap Levanter configurations with resource allocation and environment variables
- **Data processing** uses specialized configs such as `SampleCapConfig` and `WebDataCommonsStagingConfig` to control ingestion and transformation
- **Validation rules** enforce hardware compatibility, positive timeouts, and logical hierarchy constraints across all configuration groups

## Frequently Asked Questions

### How do I configure a remote inference endpoint in Marin?

Construct a `RemoteInferenceConfig` by composing `ServedModelConfig` for model weights, `VllmEngineConfig` for the runtime, and `IrisConfig` for resource placement. Set the `instances` field to specify worker count, then pass the configuration to the Iris service. All fields support validation to ensure compatible hardware and positive timeout values.

### What validation rules apply to Marin configuration options?

Marin enforces several validation constraints: `weights` paths must be non-empty in `ServedModelConfig`; `max_model_len` and `tensor_parallel_size` must be positive integers; `BrokerConfig` requires the inequality `0 < worker_timeout < lease_timeout < proxy_timeout`; and TPU configurations must specify at least one accelerator. These checks run at instantiation time due to the frozen dataclass design.

### How are training resources specified in Marin?

Training resources use the `ResourceConfig` class from `fray.types`, embedded within `TrainLmOnPodConfig` or `TrainDpoOnPodConfig`. You specify CPU, GPU, or TPU allocation through helper methods like `ResourceConfig.with_tpu()` or `ResourceConfig.with_gpu()`, passing accelerator-specific configurations such as `TpuConfig(variant="v4-8", vm_count=1)`.

### Can Marin configurations be serialized to YAML or JSON?

Yes. Because all configuration objects are standard Python dataclasses, they can be serialized to YAML or JSON using standard libraries like `dataclasses.asdict()` or Pydantic (if wrapped). This enables reproducible pipeline definitions and CLI-driven workflows where configurations are loaded from disk rather than constructed programmatically.