Configuration Options for Marin: A Complete Guide to the ML Pipeline Framework
Marin uses immutable dataclasses organized into inference, training, and data-processing groups to control every aspect of model serving and training pipelines.
Marin is a lightweight, composable framework for data ingestion, model training, and inference. All runtime behavior is governed by a hierarchy of frozen dataclasses defined in lib/marin/src/marin/, which can be instantiated via YAML/JSON or passed directly through the Python API. This guide covers the complete set of configuration options for Marin, from vLLM engine parameters to TPU resource allocation.
Inference Configuration Options
Marin’s inference stack is configured through a nested set of dataclasses that control model serving, engine selection, and request routing.
ServedModelConfig
The ServedModelConfig class in lib/marin/src/marin/inference/config.py (lines 41-50) defines how a model is exposed via an OpenAI-compatible endpoint.
Key fields include:
weights(str): Path to model artifacts (must be non-empty)revision(str | None): Specific model revision to loadapi_model(str | None): API identifier for the served modeldtype(str): Data type for weights, defaults to"bfloat16"max_model_len(int | None): Maximum sequence length (must be positive if supplied)tensor_parallel_size(int | None): Number of GPUs for tensor parallelism
VllmEngineConfig
The VllmEngineConfig class (lines 77-86) selects the vLLM runtime and controls compilation caching.
from marin.inference.config import VllmEngineConfig, VllmLauncherType, VllmSource
engine_cfg = VllmEngineConfig(
launcher=VllmLauncherType.CUDA,
source=VllmSource.MARIN_FORK,
version="0.25.1",
compilation_cache=VllmCompilationCacheMode.MANAGED,
startup_timeout_seconds=1800,
extra_args=("--enforce-eager",)
)
Validation enforces that startup_timeout_seconds must be greater than 0, and using source=VllmSource.MARIN_FORK requires launcher=VllmLauncherType.CUDA.
InferenceProxyConfig and InferenceWorkerConfig
Network behavior is controlled by InferenceProxyConfig (lines 117-124) and InferenceWorkerConfig (lines 142-149):
InferenceProxyConfigmanages the HTTP front-end with fields likeport,request_timeout_seconds(default 300.0),max_pending_requests(default 256), andresponse_fetch_batch_size(default 64)InferenceWorkerConfiglimits per-worker concurrency viamax_in_flight(default 16) andrequest_timeout_seconds(default 180.0)
All timeout and count values must be positive, and ports cannot be negative.
BrokerConfig and Iris Placement
For distributed inference, BrokerConfig (lines 154-169) multiplexes workers and proxies. It bundles worker and proxy configurations while enforcing that the broker must run on a non-preemptible resource. The config validates that timeout values satisfy the hierarchy: 0 < worker < lease < proxy.
IrisConfig (lines 92-102) handles remote inference placement through the Iris job-orchestration layer, specifying worker_resources, cache_ttl_days (default 14), and retry policies. TPU instances configured here must use single-host VMs.
RemoteInferenceConfig
The top-level RemoteInferenceConfig (lines 227-236) composes model, engine, and placement:
from marin.inference.config import (
ServedModelConfig, VllmEngineConfig, RemoteInferenceConfig, IrisConfig
)
from fray.types import ResourceConfig, GpuConfig, EnvironmentConfig
remote_cfg = RemoteInferenceConfig(
model=ServedModelConfig(
weights="gs://my-bucket/llama-7b",
api_model="llama-7b",
dtype="bfloat16"
),
engine=VllmEngineConfig(launcher=VllmLauncherType.CUDA),
iris=IrisConfig(
worker_resources=ResourceConfig.with_gpu(gpu=GpuConfig(num_gpus=1)),
worker_environment=EnvironmentConfig(),
cache_ttl_days=7
),
instances=2
)
The instances field must be positive and determines how many workers Iris spawns.
Training Configuration Options
Marin wraps the Levanter training library with configuration classes that handle resource allocation and environment setup.
TrainLmOnPodConfig and TrainDpoOnPodConfig
Defined in lib/marin/src/marin/training/training.py (lines 90-107 and 109-126), these classes manage LM and DPO training on GKE/TPU pods:
from marin.training.training import TrainLmOnPodConfig
from levanter.main.train_lm import TrainLmConfig
from fray.types import ResourceConfig, TpuConfig
pod_cfg = TrainLmOnPodConfig(
train_config=TrainLmConfig(model=model_def, trainer=trainer_cfg),
resources=ResourceConfig.with_tpu(
tpu=TpuConfig(variant="v4-8", vm_count=1)
),
output_path="gs://my-bucket/experiments/run-123",
env_vars={"WANDB_API_KEY": "..."},
auto_build_caches=False # Default: prevents accidental cache creation
)
TrainDpoOnPodConfig adds auto_num_epochs and auto_validation_runs fields that auto-compute training steps when specified.
Resource Allocation
The ResourceConfig class from fray.types (referenced in training configs) describes CPU, GPU, and TPU allocation with fields for cpu, ram, disk, gpu, tpu, and preemptible flags. Validation ensures TPU configurations include at least one accelerator.
Data Ingestion and Transformation Configuration
Marin’s Datakit and Transform sub-packages expose specialized configurations for pipeline preprocessing.
Datakit Configuration
SampleCapConfig in lib/marin/src/marin/datakit/ingestion_manifest.py (lines 71-77) controls ingestion limits:
from marin.datakit.ingestion_manifest import SampleCapConfig
cap_cfg = SampleCapConfig(
max_bytes=1e9,
max_examples=10000
)
Transform Pipeline Configuration
The transform layer includes format-specific configurations:
WebDataCommonsStagingConfig(lib/marin/src/marin/transform/structured_text/web_data_commons.py, lines 49-58): Controls staging of WebDataCommons corpus with bucket and sample cap settingsResiliparseConfig(lib/marin/src/marin/schemas/web/convert.py, lines 51-55): Configures HTML-to-Markdown conversion behaviorStackExchangeExtractionConfig(lib/marin/src/marin/transform/stackexchange/transform_stackexchange.py, lines 28-34): Governs extraction from StackExchange dumpsHfRawTextSurfaceConfig: Manages surface-level processing of HuggingFace text datasetsZeekToDolmaConfig: Converts Zeek network-security logs into the Dolma schema
How Configuration Objects Work
All Marin configuration options are implemented as frozen dataclasses (@dataclass(frozen=True)), ensuring immutability after construction. This design guarantees reproducibility and enables safe serialization across distributed workers.
The validation layer enforces sensible defaults at instantiation time. For example, BrokerConfig validates that timeout values follow a strict hierarchy, while IrisConfig ensures TPU placements use single-host VMs. Configurations can be instantiated programmatically, as shown in the examples above, or parsed from YAML/JSON files for CLI usage.
Summary
- Immutable dataclasses form the foundation of Marin configuration, located in
lib/marin/src/marin/inference/config.pyandlib/marin/src/marin/training/training.py - Inference options include
ServedModelConfigfor model artifacts,VllmEngineConfigfor runtime selection, andRemoteInferenceConfigfor orchestration - Training configs like
TrainLmOnPodConfigwrap Levanter configurations with resource allocation and environment variables - Data processing uses specialized configs such as
SampleCapConfigandWebDataCommonsStagingConfigto control ingestion and transformation - Validation rules enforce hardware compatibility, positive timeouts, and logical hierarchy constraints across all configuration groups
Frequently Asked Questions
How do I configure a remote inference endpoint in Marin?
Construct a RemoteInferenceConfig by composing ServedModelConfig for model weights, VllmEngineConfig for the runtime, and IrisConfig for resource placement. Set the instances field to specify worker count, then pass the configuration to the Iris service. All fields support validation to ensure compatible hardware and positive timeout values.
What validation rules apply to Marin configuration options?
Marin enforces several validation constraints: weights paths must be non-empty in ServedModelConfig; max_model_len and tensor_parallel_size must be positive integers; BrokerConfig requires the inequality 0 < worker_timeout < lease_timeout < proxy_timeout; and TPU configurations must specify at least one accelerator. These checks run at instantiation time due to the frozen dataclass design.
How are training resources specified in Marin?
Training resources use the ResourceConfig class from fray.types, embedded within TrainLmOnPodConfig or TrainDpoOnPodConfig. You specify CPU, GPU, or TPU allocation through helper methods like ResourceConfig.with_tpu() or ResourceConfig.with_gpu(), passing accelerator-specific configurations such as TpuConfig(variant="v4-8", vm_count=1).
Can Marin configurations be serialized to YAML or JSON?
Yes. Because all configuration objects are standard Python dataclasses, they can be serialized to YAML or JSON using standard libraries like dataclasses.asdict() or Pydantic (if wrapped). This enables reproducible pipeline definitions and CLI-driven workflows where configurations are loaded from disk rather than constructed programmatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →