Core Components of the Marin Project: A Deep Dive into the LLM Training Stack
The marin project consists of six tightly-coupled Python libraries—Levanter, Iris, Zephyr, Fray, Rigging, and Marin Core—that together provide a complete JAX-based LLM training stack covering everything from data pipeline management to distributed execution.
The marin-community/marin repository structures large-scale language model training as a modular framework. Understanding the core components of the marin project is essential for developers looking to orchestrate complex workflows, as the architecture deliberately separates concerns across specialized libraries handling optimization, resource allocation, data processing, and experiment execution.
Overview of the Core Components
The marin project organizes functionality into six specialized libraries layered to form a complete training stack. Higher-level orchestration code in Marin Core builds upon primitives supplied by Levanter, Iris, Zephyr, Fray, and Rigging, creating a dependency graph that executes data preparation, training, and checkpointing steps in the correct topological order.
The Six Core Libraries Explained
Levanter: JAX-Based Training Library
Levanter provides the optimization and model scaffolding layer. Located in lib/levanter/, this component handles optimizers, checkpoint management, and training loops. The key entry point is levanter.optim.AdamConfig, implemented in lib/levanter/optim/adam.py, which supplies the optimization algorithms used throughout the training pipeline.
Iris: Cluster-Wide Orchestration
Iris manages resource allocation and task scheduling across distributed clusters. Found in lib/iris/, this library exposes iris.cluster.ResourceConfig from lib/iris/cluster.py to define hardware constraints and scheduling policies. Iris determines the topological order of execution steps and provisions the necessary compute resources before any training begins.
Zephyr: Data Pipeline Framework
Zephyr handles data ingestion, tokenization, and sharding through a streaming abstraction. The lib/zephyr/src/zephyr/writers.py module defines writer primitives, while zephyr.worker_context.WorkerContext provides the interface for distributed data workers. This component transforms raw datasets into tokenized shards ready for training consumption.
Fray: Distributed Execution Primitives
Fray supplies the low-level distributed computing infrastructure. Located in lib/fray/, it implements actors, futures, and remote procedure calls via fray.cluster.DistributedClient (defined in lib/fray/cluster.py). Fray serves as the execution backend that Iris uses to launch remote training workers across the cluster.
Rigging: Infrastructure Utilities
Rigging provides cross-cutting concerns including configuration discovery, secret management, and logging. The lib/rigging/config/discovery.py module contains rigging.config.discovery.discover_config, which auto-detects configuration files and environment variables shared across all other marin components.
Marin Core: Experiment Definition Engine
Marin Core sits at the top of the stack, providing the lazy execution engine and step dependency graph. Implemented in lib/marin/, this component exposes marin.experiment.train.train_lm from lib/marin/experiment/train.py and the StepRunner class from lib/marin/execution/step_runner.py to define and execute training experiments without immediate evaluation.
Component Interaction and Execution Flow
The core components of the marin project follow a specific execution workflow:
- Data Preparation: Zephyr reads raw sources, applies filters, and writes tokenized shards using the writers defined in
lib/zephyr/src/zephyr/writers.py. - Training: Levanter consumes the tokenized shards under resource constraints defined by Iris, storing checkpoints via JAX-based training loops.
- Orchestration: Iris launches steps in topological order while Fray provides the remote execution backend through
DistributedClient. - Utilities: Rigging supplies configuration discovery from
lib/rigging/config/discovery.pyand logging across all phases.
Practical Implementation Examples
Tokenizing Data with Zephyr
from marin.experiment.data import tokenized
from experiments.marin_tokenizer import marin_tokenizer
# Create a lazy handle describing how to tokenize TinyStories
tinystories = tokenized(
name="tokenized/tinystories",
source="roneneldan/TinyStories",
tokenizer=marin_tokenizer,
sample_count=1_000, # cap per shard for the tutorial
)
Configuring Training with Levanter
from levanter.optim import AdamConfig
from marin.experiment.train import train_lm
from experiments.llama import llama_nano
nano_model = train_lm(
name="checkpoints/marin-nano-tinystories",
version="v1",
model=llama_nano,
optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
datasets={tinystories: 1.0}, # declare tokenized dataset as dependency
batch_size=4,
seq_len=2048,
num_train_steps=100,
)
Executing Experiments with Marin Core
from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
if __name__ == "__main__":
StepRunner().run([lower(nano_model)])
The StepRunner builds the dependency graph, requests resources from Iris, and uses Fray to launch the training worker that executes the Levanter-based training step.
Critical Source Files and APIs
The following files define the public APIs for the core components of the marin project:
lib/marin/execution/step_runner.py: Top-level driver that resolves step dependencies and triggers executionlib/marin/execution/lazy.py: Helpers for transforming lazy handles into concrete execution stepslib/levanter/optim/adam.py: Adam optimizer implementation used by training stepslib/iris/cluster.py: Resource configuration utilities and cluster managementlib/zephyr/src/zephyr/writers.py: Core writer abstractions for dataset processinglib/fray/cluster.py: Distributed client implementation for remote worker launchinglib/rigging/config/discovery.py: Configuration auto-detection and environment handling
Summary
- The marin project comprises six specialized libraries: Levanter (training), Iris (orchestration), Zephyr (data), Fray (execution), Rigging (utilities), and Marin Core (experiment definition).
- Each component exposes specific entry points such as
levanter.optim.AdamConfigandiris.cluster.ResourceConfigto configure distinct aspects of the training pipeline. - Execution follows a layered workflow where Zephyr prepares data, Levanter performs training under Iris resource management, and Fray handles remote execution, all coordinated by Marin Core's dependency graph.
- Key source files including
lib/marin/execution/step_runner.pyandlib/levanter/optim/adam.pyprovide the primary extension points for developers modifying training behavior.
Frequently Asked Questions
What is the relationship between Marin Core and Levanter?
Marin Core provides the high-level experiment definition and lazy execution graph, while Levanter supplies the underlying JAX-based training primitives including optimizers and checkpoint handling. Marin Core calls into Levanter's APIs (such as AdamConfig from lib/levanter/optim/adam.py) to execute the actual training loops on prepared data.
How does Iris differ from Fray in the marin architecture?
Iris handles high-level cluster-wide resource allocation and scheduling decisions, exposing ResourceConfig in lib/iris/cluster.py to define hardware constraints. Fray provides the low-level distributed execution primitives including remote procedure calls and actor management through DistributedClient in lib/fray/cluster.py, which Iris uses as its execution backend.
Where does data preprocessing occur in the marin project?
Data preprocessing and tokenization occur in the Zephyr component, specifically utilizing writer abstractions defined in lib/zephyr/src/zephyr/writers.py. The tokenized() helper function from marin.experiment.data creates lazy dataset handles that describe how raw sources transform into training-ready shards.
Which file should I modify to change how marin discovers configuration files?
Configuration discovery logic resides in lib/rigging/config/discovery.py, which implements rigging.config.discovery.discover_config. This module auto-detects configuration files and environment variables across all marin components, making it the primary extension point for customizing configuration loading behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →