Core Components of the Marin Project: A Deep Dive into the LLM Training Stack

The marin project consists of six tightly-coupled Python libraries—Levanter, Iris, Zephyr, Fray, Rigging, and Marin Core—that together provide a complete JAX-based LLM training stack covering everything from data pipeline management to distributed execution.

The marin-community/marin repository structures large-scale language model training as a modular framework. Understanding the core components of the marin project is essential for developers looking to orchestrate complex workflows, as the architecture deliberately separates concerns across specialized libraries handling optimization, resource allocation, data processing, and experiment execution.

Overview of the Core Components

The marin project organizes functionality into six specialized libraries layered to form a complete training stack. Higher-level orchestration code in Marin Core builds upon primitives supplied by Levanter, Iris, Zephyr, Fray, and Rigging, creating a dependency graph that executes data preparation, training, and checkpointing steps in the correct topological order.

The Six Core Libraries Explained

Levanter: JAX-Based Training Library

Levanter provides the optimization and model scaffolding layer. Located in lib/levanter/, this component handles optimizers, checkpoint management, and training loops. The key entry point is levanter.optim.AdamConfig, implemented in lib/levanter/optim/adam.py, which supplies the optimization algorithms used throughout the training pipeline.

Iris: Cluster-Wide Orchestration

Iris manages resource allocation and task scheduling across distributed clusters. Found in lib/iris/, this library exposes iris.cluster.ResourceConfig from lib/iris/cluster.py to define hardware constraints and scheduling policies. Iris determines the topological order of execution steps and provisions the necessary compute resources before any training begins.

Zephyr: Data Pipeline Framework

Zephyr handles data ingestion, tokenization, and sharding through a streaming abstraction. The lib/zephyr/src/zephyr/writers.py module defines writer primitives, while zephyr.worker_context.WorkerContext provides the interface for distributed data workers. This component transforms raw datasets into tokenized shards ready for training consumption.

Fray: Distributed Execution Primitives

Fray supplies the low-level distributed computing infrastructure. Located in lib/fray/, it implements actors, futures, and remote procedure calls via fray.cluster.DistributedClient (defined in lib/fray/cluster.py). Fray serves as the execution backend that Iris uses to launch remote training workers across the cluster.

Rigging: Infrastructure Utilities

Rigging provides cross-cutting concerns including configuration discovery, secret management, and logging. The lib/rigging/config/discovery.py module contains rigging.config.discovery.discover_config, which auto-detects configuration files and environment variables shared across all other marin components.

Marin Core: Experiment Definition Engine

Marin Core sits at the top of the stack, providing the lazy execution engine and step dependency graph. Implemented in lib/marin/, this component exposes marin.experiment.train.train_lm from lib/marin/experiment/train.py and the StepRunner class from lib/marin/execution/step_runner.py to define and execute training experiments without immediate evaluation.

Component Interaction and Execution Flow

The core components of the marin project follow a specific execution workflow:

  1. Data Preparation: Zephyr reads raw sources, applies filters, and writes tokenized shards using the writers defined in lib/zephyr/src/zephyr/writers.py.
  2. Training: Levanter consumes the tokenized shards under resource constraints defined by Iris, storing checkpoints via JAX-based training loops.
  3. Orchestration: Iris launches steps in topological order while Fray provides the remote execution backend through DistributedClient.
  4. Utilities: Rigging supplies configuration discovery from lib/rigging/config/discovery.py and logging across all phases.

Practical Implementation Examples

Tokenizing Data with Zephyr

from marin.experiment.data import tokenized
from experiments.marin_tokenizer import marin_tokenizer

# Create a lazy handle describing how to tokenize TinyStories

tinystories = tokenized(
    name="tokenized/tinystories",
    source="roneneldan/TinyStories",
    tokenizer=marin_tokenizer,
    sample_count=1_000,  # cap per shard for the tutorial

)

Configuring Training with Levanter

from levanter.optim import AdamConfig
from marin.experiment.train import train_lm
from experiments.llama import llama_nano

nano_model = train_lm(
    name="checkpoints/marin-nano-tinystories",
    version="v1",
    model=llama_nano,
    optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
    datasets={tinystories: 1.0},  # declare tokenized dataset as dependency

    batch_size=4,
    seq_len=2048,
    num_train_steps=100,
)

Executing Experiments with Marin Core

from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner

if __name__ == "__main__":
    StepRunner().run([lower(nano_model)])

The StepRunner builds the dependency graph, requests resources from Iris, and uses Fray to launch the training worker that executes the Levanter-based training step.

Critical Source Files and APIs

The following files define the public APIs for the core components of the marin project:

Summary

  • The marin project comprises six specialized libraries: Levanter (training), Iris (orchestration), Zephyr (data), Fray (execution), Rigging (utilities), and Marin Core (experiment definition).
  • Each component exposes specific entry points such as levanter.optim.AdamConfig and iris.cluster.ResourceConfig to configure distinct aspects of the training pipeline.
  • Execution follows a layered workflow where Zephyr prepares data, Levanter performs training under Iris resource management, and Fray handles remote execution, all coordinated by Marin Core's dependency graph.
  • Key source files including lib/marin/execution/step_runner.py and lib/levanter/optim/adam.py provide the primary extension points for developers modifying training behavior.

Frequently Asked Questions

What is the relationship between Marin Core and Levanter?

Marin Core provides the high-level experiment definition and lazy execution graph, while Levanter supplies the underlying JAX-based training primitives including optimizers and checkpoint handling. Marin Core calls into Levanter's APIs (such as AdamConfig from lib/levanter/optim/adam.py) to execute the actual training loops on prepared data.

How does Iris differ from Fray in the marin architecture?

Iris handles high-level cluster-wide resource allocation and scheduling decisions, exposing ResourceConfig in lib/iris/cluster.py to define hardware constraints. Fray provides the low-level distributed execution primitives including remote procedure calls and actor management through DistributedClient in lib/fray/cluster.py, which Iris uses as its execution backend.

Where does data preprocessing occur in the marin project?

Data preprocessing and tokenization occur in the Zephyr component, specifically utilizing writer abstractions defined in lib/zephyr/src/zephyr/writers.py. The tokenized() helper function from marin.experiment.data creates lazy dataset handles that describe how raw sources transform into training-ready shards.

Which file should I modify to change how marin discovers configuration files?

Configuration discovery logic resides in lib/rigging/config/discovery.py, which implements rigging.config.discovery.discover_config. This module auto-detects configuration files and environment variables across all marin components, making it the primary extension point for customizing configuration loading behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →