# Marin Project Structure: A Deep Dive into the Monorepo Architecture

> Explore the marin project structure, a Python monorepo housing libraries, experiments, infra, and docs. Discover its core data processing and distributed utilities.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-08-27

---

**The marin project is organized as a Python-centric monorepo containing tightly-coupled libraries (`lib/`), experiment scripts (`experiments/`), infrastructure definitions (`infra/`), and comprehensive documentation (`docs/`), with core data processing handled by the Zephyr lazy dataset engine and distributed utilities provided by the Rigging library.**

The marin-community/marin repository follows a layered architecture designed for large-scale machine learning pipelines. Understanding the marin project structure is essential for contributors working with its declarative data processing engine, distributed execution primitives, or pipeline orchestration tools.

## Top-Level Directory Layout

The repository root divides functionality into seven primary directories, each serving a distinct purpose in the ML lifecycle.

- **`lib/`** – Reusable Python libraries powering Marin’s pipeline, dataset handling, and distributed execution.
- **`experiments/`** – Example scripts demonstrating model training, tokenization, and scaling-law sweeps.
- **`docs/`** – User-facing knowledge base including installation guides and design overviews.
- **`infra/`** – Infrastructure-as-code (Pulumi) and ancillary services like the XProf profiling server.
- **`tests/`** – Automated verification split by sub-project (e.g., `tests/vllm/`, `tests/transform/`).
- **`config/`** – Default YAML configuration files driving pipelines and external tools.

Key repository metadata files sit at the root: [`AGENTS.md`](https://github.com/marin-community/marin/blob/main/AGENTS.md) defines global conventions, [`README.md`](https://github.com/marin-community/marin/blob/main/README.md) provides the high-level overview, and [`CONTRIBUTING.md`](https://github.com/marin-community/marin/blob/main/CONTRIBUTING.md) outlines the contribution workflow.

## Core Libraries in `lib/`

The `lib/` directory contains the foundational components of the marin project structure, implementing a three-layer architecture where Rigging provides low-level primitives, Zephyr builds the dataset engine on top, and Marin orchestrates the pipeline.

### Zephyr: Lazy Dataset Processing

Located in `lib/zephyr/`, Zephyr implements a declarative, actor-based pipeline for large-scale dataset manipulation. According to [`lib/zephyr/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/zephyr/AGENTS.md), all data flow between stages is **filesystem-backed**, with each stage writing output to a `PickleDiskChunk` on disk to minimize memory pressure.

Core implementation files include:

- **[`src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/dataset.py)** – Defines the `Dataset` class and transformation primitives (`group_by`, `deduplicate`, etc.).
- **[`src/zephyr/context.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/context.py)** – Manages `ZephyrContext` and the lifecycle of worker pools.
- **[`src/zephyr/coordinator.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/coordinator.py)** – Orchestrates task scheduling, counters, and result aggregation.
- **[`src/zephyr/worker.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/worker.py)** – Implements persistent workers that pull tasks, heartbeat, and report results.
- **[`src/zephyr/plan.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/plan.py)** – Generates physical execution plans and performs operation fusion.
- **[`src/zephyr/readers.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/readers.py) & [`src/zephyr/writers.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/writers.py)** – IO adapters for JSONL, Parquet, and Vortex formats.
- **[`src/zephyr/shuffle.py`](https://github.com/marin-community/marin/blob/main/src/zephyr/shuffle.py)** – Scatter-merge shuffle implementation used across pipelines.

Workers communicate with the coordinator via RPCs that must block on `.result()` calls to preserve ordering and avoid race conditions.

### Rigging: Distributed System Utilities

The `lib/rigging/` library supplies reusable building blocks for distributed execution and observability, validated by tests in `lib/rigging/tests/` such as [`test_distributed_lock.py`](https://github.com/marin-community/marin/blob/main/test_distributed_lock.py).

Key submodules include:

- **[`rigging/tunnel.py`](https://github.com/marin-community/marin/blob/main/rigging/tunnel.py)** – Secure RPC tunnelling between processes.
- **[`rigging/token_authority.py`](https://github.com/marin-community/marin/blob/main/rigging/token_authority.py)** – Handles secure token issuance for service-to-service authentication.
- **`rigging/telemetry/`** – Prometheus exporters and custom probes (NVIDIA, NCCL).
- **[`rigging/testing.py`](https://github.com/marin-community/marin/blob/main/rigging/testing.py)** – Test helpers for simulated clusters.
- **[`rigging/server_auth.py`](https://github.com/marin-community/marin/blob/main/rigging/server_auth.py)** – Server-side authentication middleware.

### Marin: Pipeline Orchestration

While the top-level `marin` package source is largely vendored in `lib/marin/`, the directory enforces strict conventions documented in [`lib/marin/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/marin/AGENTS.md):

- Use **`fsspec.open`** for all filesystem access to maintain cloud agnosticism.
- Stream large GCS objects rather than copying them locally.
- Reference pipeline steps instead of hard-coding bucket paths.

## Experiments and Tutorials

The `experiments/` directory contains end-to-end examples demonstrating typical workflows. The file [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py) showcases tokenizing a small dataset, building a `StepRunner`, and launching a training step using Zephyr for tokenization and Marin for step definition.

Additional scripts handle scaling-law sweeps that drive large-scale runs on TPU pods, leveraging the same pipeline abstractions found in the core libraries.

## Infrastructure and Configuration

### Infrastructure Definitions

The `infra/` directory houses Pulumi configurations for auxiliary services. The XProf profiling server at [`infra/xprof/server.py`](https://github.com/marin-community/marin/blob/main/infra/xprof/server.py) enables reproducible deployment of services required during large-scale training runs.

### Configuration Files

Default pipeline configurations reside in [`config/marin.yaml`](https://github.com/marin-community/marin/blob/main/config/marin.yaml), which drives experiments and external tools without requiring hard-coded parameters in the source.

## Testing Strategy

The marin project structure splits testing across domain-specific directories to ensure comprehensive coverage of distributed components:

- **`tests/vllm/`** – Integration tests for the VLLM inference backend, including [`tests/vllm/test_llm_inference.py`](https://github.com/marin-community/marin/blob/main/tests/vllm/test_llm_inference.py).
- **`tests/transform/`** – Unit tests for data-conversion utilities (HTML, HuggingFace datasets).
- **Library-specific tests** – Co-located with source code (e.g., `lib/zephyr/tests/`, `lib/rigging/tests/`).

Execute the full test suite using the project-wide pre-commit wrapper:

```bash
./infra/pre-commit.py --all-files --fix
uv run --no-project infra/ci/run_tests.py

```

## Documentation Organization

User-facing documentation lives under `docs/` with tutorials organized by complexity:

- **[`docs/tutorials/installation.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/installation.md)** – Environment setup and dependency installation.
- **[`docs/tutorials/first-experiment.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/first-experiment.md)** – Step-by-step guide mirroring [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py).
- **[`docs/tutorials/train-an-lm.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/train-an-lm.md)** – Documentation for the "Delphi" scaling recipe.

## Summary

- The marin project structure uses a **monorepo layout** with clear separation between libraries (`lib/`), experiments (`experiments/`), and infrastructure (`infra/`).
- **Zephyr** provides the lazy dataset processing engine with filesystem-backed dataflow, while **Rigging** supplies distributed primitives like tunneling and telemetry.
- **Pipeline orchestration** follows strict conventions using `fsspec.open` and step-based configuration rather than hard-coded paths.
- **Testing** is partitioned by concern (VLLM, transforms) alongside library-specific unit tests.
- **Documentation** and **configuration** are first-class citizens, with YAML configs in `config/` and tutorials in `docs/tutorials/`.

## Frequently Asked Questions

### What is the purpose of the `lib/` directory in the marin repository?

The `lib/` directory contains the reusable Python libraries that constitute the core of the marin project structure. It houses Zephyr for lazy dataset processing, Rigging for distributed system utilities, and the vendored Marin orchestration layer. This centralized approach allows tight integration between the dataset engine and distributed execution primitives while maintaining modular boundaries.

### How does Zephyr handle memory management during large-scale data processing?

According to [`lib/zephyr/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/zephyr/AGENTS.md), Zephyr uses a **filesystem-backed** architecture where each pipeline stage writes its output to a `PickleDiskChunk` on disk rather than keeping data in memory. The next stage reads from these chunks, minimizing in-memory pressure across the declarative pipeline. Workers coordinate via RPCs that block on `.result()` calls to maintain execution ordering without consuming excessive RAM.

### Where are the end-to-end training examples located?

Complete training examples reside in `experiments/tutorials/`, specifically [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py). This script demonstrates tokenizing a dataset with Zephyr, constructing a `StepRunner`, and executing a training step. The `docs/tutorials/` directory contains complementary written guides such as [`first-experiment.md`](https://github.com/marin-community/marin/blob/main/first-experiment.md) that walk through the same workflow step-by-step.

### How is authentication handled between distributed services in marin?

Authentication is managed by the Rigging library in `lib/rigging/`. The [`rigging/token_authority.py`](https://github.com/marin-community/marin/blob/main/rigging/token_authority.py) module handles secure token issuance for service-to-service authentication, while [`rigging/server_auth.py`](https://github.com/marin-community/marin/blob/main/rigging/server_auth.py) provides server-side authentication middleware. These components are tested in [`lib/rigging/tests/test_distributed_lock.py`](https://github.com/marin-community/marin/blob/main/lib/rigging/tests/test_distributed_lock.py) and related files to ensure secure communication across distributed training clusters.