# Main Modules and Packages in the Marin Repository: A Complete Architecture Guide

> Explore the main modules and packages in the Marin repository. Discover how these nine Python libraries form a layered stack for large-scale language model research from data ingestion to distributed training.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: architecture
- Published: 2026-08-29

---

**The Marin repository organizes its functionality into nine specialized Python libraries located under the top-level `lib/` directory, forming a layered dependency stack that enables large-scale language model research from data ingestion through distributed training.**

The `marin-community/marin` codebase is structured as a collection of loosely-coupled packages designed to handle specific stages of the ML lifecycle. Understanding the main modules or packages in Marin is essential for anyone contributing to or building upon this platform, as the architecture deliberately enforces unidirectional dependencies where higher-level libraries import only from lower-level ones.

## Core Library Architecture and Design Philosophy

The Marin repository follows a strict hierarchical dependency model to maintain clean separation of concerns. The dependency flow runs from low-level data utilities upward to high-level orchestration: **`iris`** → **`levanter`** → **`zephyr`** → **`marin`**.

This architecture ensures that foundational utilities remain independent of business logic, while the top-level `marin` package can compose functionality from all underlying layers. All core libraries reside under the `lib/` directory, with each package containing its source code in `lib/<package>/src/<package>/`.

## The Nine Main Packages in Marin

### Zephyr: Low-Level Data Pipeline Utilities

**Zephyr** provides the foundational data-processing layer, implementing readers, writers, workers, and I/O optimizations. Located at `lib/zephyr`, this package handles the initial data ingestion and transformation stages.

Key implementation files include [`lib/zephyr/src/zephyr/worker.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/worker.py), which manages parallel data loading workers. Typical usage involves reading Parquet datasets directly from cloud storage:

```python
from lib.zephyr import ParquetReader

reader = ParquetReader(path="gs://my-bucket/dataset.parquet")
for record in reader:
    # process each record …

    pass

```

### Levanter: JAX-Based Training Framework

**Levanter** constitutes the computational engine of the stack, built on JAX for high-performance model training. Found at `lib/levanter`, it contains model definitions, trainer implementations, and optimizer configurations.

The primary entry point is [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py), which defines the core training loops and checkpoint management. Levanter consumes data processed by Zephyr and provides the `Trainer` class used by higher-level orchestration.

### Iris: Job Orchestration and Scheduling

**Iris** manages distributed execution across compute clusters, handling job scheduling and resource allocation. Located at `lib/iris`, this package bridges the gap between training logic and hardware infrastructure.

The scheduler implementation in [`lib/iris/src/iris/job_scheduler.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/job_scheduler.py) provides the `JobScheduler` class for submitting and monitoring training workloads on TPU or GPU clusters.

### Marin: High-Level Pipeline Orchestration

The **`marin`** package itself (located at `lib/marin`) provides the highest-level API, gluing together Iris, Levanter, and Zephyr for end-to-end experiment management. It defines abstractions for experiments, configurations, and workflow composition.

The core abstraction resides in [`lib/marin/src/marin/experiment.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment.py), which enables researchers to define complete training pipelines:

```python
from lib.marin import Experiment, Config
from lib.levanter import Trainer
from lib.iris import JobScheduler

cfg = Config(
    model="gpt2-small",
    optimizer="adamw",
    dataset="openwebtext",
    batch_size=128,
)

trainer = Trainer(cfg)
scheduler = JobScheduler(cluster="tpu-v4")
experiment = Experiment(name="gpt2_small_demo", trainer=trainer)

scheduler.submit(experiment)

```

### Fray: Distributed Execution Runtime

**Fray** implements the actor-model runtime for distributed computation, managing actors, futures, and client-server communication. Located at `lib/fray`, it provides the underlying execution primitives used by Iris for cluster management.

Key components are defined in [`lib/fray/src/fray/actor.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/actor.py), which implements the actor-based concurrency model.

### Rigging: Infrastructure and Utilities

**Rigging** supplies cross-cutting infrastructure concerns including authentication, secrets management, telemetry, and storage abstractions. Located at `lib/rigging`, it provides shared utilities used across the entire stack.

The secrets management system in [`lib/rigging/src/rigging/secrets.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/secrets.py) exemplifies this utility:

```python
from lib.rigging.secrets import SecretStore

store = SecretStore()
api_key = store.get("openai_api_key")

```

### Ducky: RPC and Tunneling Framework

**Ducky** is a lightweight RPC and tunneling framework used primarily by Rigging and Fray for inter-service communication. Located at `lib/ducky`, it implements the networking layer required for distributed coordination.

Core tunneling logic resides in [`lib/ducky/src/ducky/tunnel.py`](https://github.com/marin-community/marin/blob/main/lib/ducky/src/ducky/tunnel.py).

### Dupekit: Data Deduplication Utilities

**Dupekit** provides specialized algorithms for detecting and removing duplicate content in training datasets. Located at `lib/dupekit`, it implements hashing strategies including MinHash and Bloom filters for efficient deduplication at scale.

The MinHash implementation in [`lib/dupekit/src/dupekit/minhash.py`](https://github.com/marin-community/marin/blob/main/lib/dupekit/src/dupekit/minhash.py) serves as the primary interface for near-duplicate detection in data preprocessing pipelines.

### Finestore: Persistent Artifact Storage

**Finestore** manages versioned storage for datasets and model artifacts, providing filesystem abstractions, caching layers, and schema management. Located at `lib/finestore`, it ensures reproducibility by tracking data lineage and model versions.

The storage interface is defined in [`lib/finestore/src/finestore/store.py`](https://github.com/marin-community/marin/blob/main/lib/finestore/src/finestore/store.py).

## Additional Repository Components

Beyond the core `lib/` packages, the Marin repository includes several supporting directories:

- **`experiments/`** – Example training scripts and tutorials demonstrating package integration patterns
- **`infra/`** – Pulumi-based Infrastructure-as-Code definitions for deploying TPU clusters and GCP resources
- **`docs/`** – MkDocs-based documentation covering installation guides and architectural overviews

## Summary

The Marin repository organizes its functionality into nine distinct Python libraries under `lib/`, each addressing specific requirements of large-scale ML research:

- **Zephyr** handles low-level data I/O and pipeline utilities at [`lib/zephyr/src/zephyr/worker.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/worker.py)
- **Levanter** provides the JAX-based training framework with entry points in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py)
- **Iris** manages cluster job orchestration through [`lib/iris/src/iris/job_scheduler.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/job_scheduler.py)
- **Marin** offers high-level experiment composition via [`lib/marin/src/marin/experiment.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment.py)
- **Fray** implements the distributed execution runtime in [`lib/fray/src/fray/actor.py`](https://github.com/marin-community/marin/blob/main/lib/fray/src/fray/actor.py)
- **Rigging** supplies authentication and secrets handling through [`lib/rigging/src/rigging/secrets.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/src/rigging/secrets.py)
- **Ducky** enables lightweight RPC communication via [`lib/ducky/src/ducky/tunnel.py`](https://github.com/marin-community/marin/blob/main/lib/ducky/src/ducky/tunnel.py)
- **Dupekit** performs dataset deduplication using [`lib/dupekit/src/dupekit/minhash.py`](https://github.com/marin-community/marin/blob/main/lib/dupekit/src/dupekit/minhash.py)
- **Finestore** manages versioned storage through [`lib/finestore/src/finestore/store.py`](https://github.com/marin-community/marin/blob/main/lib/finestore/src/finestore/store.py)

This layered architecture, with its strict dependency hierarchy from Zephyr up to Marin, enables scalable, reproducible language model research while maintaining clear separation of concerns between data processing, training computation, and orchestration logic.

## Frequently Asked Questions

### What is the purpose of the Levanter package in Marin?

**Levanter serves as the JAX-based training framework within Marin, handling model definitions, optimization loops, and distributed training logic.** It consumes preprocessed data from Zephyr and exposes trainer implementations that Iris schedules across compute clusters. The core training logic resides in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py).

### How does the Marin package differ from the Iris package?

**While Iris focuses specifically on job scheduling and cluster resource management, Marin provides high-level pipeline orchestration that coordinates multiple subsystems.** Marin imports from both Iris (for scheduling) and Levanter (for training) to compose end-to-end experiments, whereas Iris operates independently at the infrastructure layer managing job lifecycle on TPU/GPU clusters.

### Where are data deduplication utilities located in the Marin codebase?

**Data deduplication functionality resides in the Dupekit package at `lib/dupekit`.** This module implements MinHash algorithms and Bloom filters specifically designed for preprocessing large training datasets to remove duplicate content before ingestion into the training pipeline, with core algorithms in [`lib/dupekit/src/dupekit/minhash.py`](https://github.com/marin-community/marin/blob/main/lib/dupekit/src/dupekit/minhash.py).

### What file path contains the core Experiment class for defining training jobs?

**The `Experiment` class is defined in [`lib/marin/src/marin/experiment.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment.py).** This class serves as the primary abstraction for configuring and launching complete training workflows, integrating configurations from Levanter with scheduling capabilities from Iris to provide the high-level API for the Marin platform.