Configuration Options for Marin: A Complete Guide to the ML Pipeline Framework

Marin uses immutable dataclasses organized into inference, training, and data-processing groups to control every aspect of model serving and training pipelines.

Marin is a lightweight, composable framework for data ingestion, model training, and inference. All runtime behavior is governed by a hierarchy of frozen dataclasses defined in lib/marin/src/marin/, which can be instantiated via YAML/JSON or passed directly through the Python API. This guide covers the complete set of configuration options for Marin, from vLLM engine parameters to TPU resource allocation.

Inference Configuration Options

Marin’s inference stack is configured through a nested set of dataclasses that control model serving, engine selection, and request routing.

ServedModelConfig

The ServedModelConfig class in lib/marin/src/marin/inference/config.py (lines 41-50) defines how a model is exposed via an OpenAI-compatible endpoint.

Key fields include:

  • weights (str): Path to model artifacts (must be non-empty)
  • revision (str | None): Specific model revision to load
  • api_model (str | None): API identifier for the served model
  • dtype (str): Data type for weights, defaults to "bfloat16"
  • max_model_len (int | None): Maximum sequence length (must be positive if supplied)
  • tensor_parallel_size (int | None): Number of GPUs for tensor parallelism

VllmEngineConfig

The VllmEngineConfig class (lines 77-86) selects the vLLM runtime and controls compilation caching.

from marin.inference.config import VllmEngineConfig, VllmLauncherType, VllmSource

engine_cfg = VllmEngineConfig(
    launcher=VllmLauncherType.CUDA,
    source=VllmSource.MARIN_FORK,
    version="0.25.1",
    compilation_cache=VllmCompilationCacheMode.MANAGED,
    startup_timeout_seconds=1800,
    extra_args=("--enforce-eager",)
)

Validation enforces that startup_timeout_seconds must be greater than 0, and using source=VllmSource.MARIN_FORK requires launcher=VllmLauncherType.CUDA.

InferenceProxyConfig and InferenceWorkerConfig

Network behavior is controlled by InferenceProxyConfig (lines 117-124) and InferenceWorkerConfig (lines 142-149):

  • InferenceProxyConfig manages the HTTP front-end with fields like port, request_timeout_seconds (default 300.0), max_pending_requests (default 256), and response_fetch_batch_size (default 64)
  • InferenceWorkerConfig limits per-worker concurrency via max_in_flight (default 16) and request_timeout_seconds (default 180.0)

All timeout and count values must be positive, and ports cannot be negative.

BrokerConfig and Iris Placement

For distributed inference, BrokerConfig (lines 154-169) multiplexes workers and proxies. It bundles worker and proxy configurations while enforcing that the broker must run on a non-preemptible resource. The config validates that timeout values satisfy the hierarchy: 0 < worker < lease < proxy.

IrisConfig (lines 92-102) handles remote inference placement through the Iris job-orchestration layer, specifying worker_resources, cache_ttl_days (default 14), and retry policies. TPU instances configured here must use single-host VMs.

RemoteInferenceConfig

The top-level RemoteInferenceConfig (lines 227-236) composes model, engine, and placement:

from marin.inference.config import (
    ServedModelConfig, VllmEngineConfig, RemoteInferenceConfig, IrisConfig
)
from fray.types import ResourceConfig, GpuConfig, EnvironmentConfig

remote_cfg = RemoteInferenceConfig(
    model=ServedModelConfig(
        weights="gs://my-bucket/llama-7b",
        api_model="llama-7b",
        dtype="bfloat16"
    ),
    engine=VllmEngineConfig(launcher=VllmLauncherType.CUDA),
    iris=IrisConfig(
        worker_resources=ResourceConfig.with_gpu(gpu=GpuConfig(num_gpus=1)),
        worker_environment=EnvironmentConfig(),
        cache_ttl_days=7
    ),
    instances=2
)

The instances field must be positive and determines how many workers Iris spawns.

Training Configuration Options

Marin wraps the Levanter training library with configuration classes that handle resource allocation and environment setup.

TrainLmOnPodConfig and TrainDpoOnPodConfig

Defined in lib/marin/src/marin/training/training.py (lines 90-107 and 109-126), these classes manage LM and DPO training on GKE/TPU pods:

from marin.training.training import TrainLmOnPodConfig
from levanter.main.train_lm import TrainLmConfig
from fray.types import ResourceConfig, TpuConfig

pod_cfg = TrainLmOnPodConfig(
    train_config=TrainLmConfig(model=model_def, trainer=trainer_cfg),
    resources=ResourceConfig.with_tpu(
        tpu=TpuConfig(variant="v4-8", vm_count=1)
    ),
    output_path="gs://my-bucket/experiments/run-123",
    env_vars={"WANDB_API_KEY": "..."},
    auto_build_caches=False  # Default: prevents accidental cache creation

)

TrainDpoOnPodConfig adds auto_num_epochs and auto_validation_runs fields that auto-compute training steps when specified.

Resource Allocation

The ResourceConfig class from fray.types (referenced in training configs) describes CPU, GPU, and TPU allocation with fields for cpu, ram, disk, gpu, tpu, and preemptible flags. Validation ensures TPU configurations include at least one accelerator.

Data Ingestion and Transformation Configuration

Marin’s Datakit and Transform sub-packages expose specialized configurations for pipeline preprocessing.

Datakit Configuration

SampleCapConfig in lib/marin/src/marin/datakit/ingestion_manifest.py (lines 71-77) controls ingestion limits:

from marin.datakit.ingestion_manifest import SampleCapConfig

cap_cfg = SampleCapConfig(
    max_bytes=1e9,
    max_examples=10000
)

Transform Pipeline Configuration

The transform layer includes format-specific configurations:

How Configuration Objects Work

All Marin configuration options are implemented as frozen dataclasses (@dataclass(frozen=True)), ensuring immutability after construction. This design guarantees reproducibility and enables safe serialization across distributed workers.

The validation layer enforces sensible defaults at instantiation time. For example, BrokerConfig validates that timeout values follow a strict hierarchy, while IrisConfig ensures TPU placements use single-host VMs. Configurations can be instantiated programmatically, as shown in the examples above, or parsed from YAML/JSON files for CLI usage.

Summary

  • Immutable dataclasses form the foundation of Marin configuration, located in lib/marin/src/marin/inference/config.py and lib/marin/src/marin/training/training.py
  • Inference options include ServedModelConfig for model artifacts, VllmEngineConfig for runtime selection, and RemoteInferenceConfig for orchestration
  • Training configs like TrainLmOnPodConfig wrap Levanter configurations with resource allocation and environment variables
  • Data processing uses specialized configs such as SampleCapConfig and WebDataCommonsStagingConfig to control ingestion and transformation
  • Validation rules enforce hardware compatibility, positive timeouts, and logical hierarchy constraints across all configuration groups

Frequently Asked Questions

How do I configure a remote inference endpoint in Marin?

Construct a RemoteInferenceConfig by composing ServedModelConfig for model weights, VllmEngineConfig for the runtime, and IrisConfig for resource placement. Set the instances field to specify worker count, then pass the configuration to the Iris service. All fields support validation to ensure compatible hardware and positive timeout values.

What validation rules apply to Marin configuration options?

Marin enforces several validation constraints: weights paths must be non-empty in ServedModelConfig; max_model_len and tensor_parallel_size must be positive integers; BrokerConfig requires the inequality 0 < worker_timeout < lease_timeout < proxy_timeout; and TPU configurations must specify at least one accelerator. These checks run at instantiation time due to the frozen dataclass design.

How are training resources specified in Marin?

Training resources use the ResourceConfig class from fray.types, embedded within TrainLmOnPodConfig or TrainDpoOnPodConfig. You specify CPU, GPU, or TPU allocation through helper methods like ResourceConfig.with_tpu() or ResourceConfig.with_gpu(), passing accelerator-specific configurations such as TpuConfig(variant="v4-8", vm_count=1).

Can Marin configurations be serialized to YAML or JSON?

Yes. Because all configuration objects are standard Python dataclasses, they can be serialized to YAML or JSON using standard libraries like dataclasses.asdict() or Pydantic (if wrapped). This enables reproducible pipeline definitions and CLI-driven workflows where configurations are loaded from disk rather than constructed programmatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →