# Marin Repository Directory Structure: A Complete Guide to the Multi-Layered Python Codebase

> Explore the marin repository directory structure. Understand the layered Python codebase organizing core libraries, infrastructure, experiments, and docs for efficient machine learning development.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-08-29

---

**The marin repository organizes its JAX-based machine learning codebase into a strictly layered hierarchy with core libraries under `lib/`, infrastructure definitions in `infra/`, runnable experiments in `experiments/`, and documentation in `docs/`, enforcing clean dependency directions where lower-level packages never import from higher-level ones.**

The marin repository maintained by marin-community implements a modular architecture designed for large-scale training pipelines. Understanding the marin repository directory structure is essential for navigating the six distinct core packages, Pulumi-based deployment stacks, and agent-specific tooling that collectively power this distributed ML framework.

## Top-Level Directory Layout

The root of the repository contains configuration files, documentation entry points, and seven primary directories that separate concerns by function:

- **`lib/`** – Houses six core Python packages (levanter, marin, iris, zephyr, fray, rigging) that form a layered execution stack.
- **`infra/`** – Contains Pulumi-based infrastructure definitions for TPU profiling, CI management, and deployment patterns.
- **`experiments/`** – Holds runnable tutorials, tokenization scripts, and supervised fine-tuning (SFT) pipelines.
- **`docs/`** – Source files for the MkDocs-generated documentation site, including tutorials and system prompts.
- **`config/`** – Default YAML and TOML configuration files for various environments (e.g., [`config/marin.yaml`](https://github.com/marin-community/marin/blob/main/config/marin.yaml), [`config/coreweave.yaml`](https://github.com/marin-community/marin/blob/main/config/coreweave.yaml)).
- **`tests/`** – Integration tests for VLLM interfaces and cross-package validation suites.
- **`.github/`** – CI workflows, issue templates, and automation scripts.

Additional hidden directories support repository automation: `.agents/` stores agent skills and writing-style guidelines, while `.claude/` contains internal tooling settings. Root configuration files include [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) (build system and dependencies), [`mkdocs.yml`](https://github.com/marin-community/marin/blob/main/mkdocs.yml) (documentation site configuration), [`AGENTS.md`](https://github.com/marin-community/marin/blob/main/AGENTS.md) (agent development policies), and [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md) (testing conventions).

## Core Libraries in lib/

The `lib/` directory implements a dependency-directed architecture where packages at lower layers provide primitives to higher layers. Each package contains a `src/` directory for production code, a nested `tests/` folder, and an [`AGENTS.md`](https://github.com/marin-community/marin/blob/main/AGENTS.md) file documenting package-specific guidelines.

### levanter

Located at `lib/levanter/`, this package provides JAX-based training primitives including model definitions, optimizer configurations, and checkpointing logic. Code in this layer handles the fundamental mathematics of distributed training without awareness of pipeline orchestration.

### marin

The high-level pipeline framework at `lib/marin/` composes training steps into executable experiments. The critical subdirectory `lib/marin/src/marin/experiment/` defines core abstractions for experiment steps and data handling. This layer imports from levanter but remains isolated from infrastructure concerns.

### iris

The job orchestration layer at `lib/iris/` manages cluster scheduling and resource allocation. It translates high-level experiment requirements from marin into concrete execution plans without containing training logic itself.

### zephyr

Dataset processing utilities live in `lib/zephyr/`, providing shuffling, I/O optimization, and format readers such as `ParquetReader`. This layer sits below the training stack but above raw storage access.

### fray

Distributed execution primitives reside in `lib/fray/`, implementing communication patterns used by higher-level packages. Fray abstracts the mechanics of multi-node synchronization.

### rigging

Low-level services including storage backends, authentication, and telemetry are defined in `lib/rigging/`. This foundational layer has no dependencies on other `lib/` packages.

## Infrastructure with infra/

The `infra/` directory contains Pulumi TypeScript stacks and helper scripts for deploying and monitoring training infrastructure:

- **`infra/xprof/`** – Pulumi stack and server code for the XProf TPU profiling service.
- **`infra/tpu-ci/`** – Scripts to launch and manage TPU CI pods for continuous integration.
- **`infra/ducky/`** – Minimal Pulumi example used exclusively for CI validation of infrastructure templates.
- **[`infra/pulumi.md`](https://github.com/marin-community/marin/blob/main/infra/pulumi.md)** – Documentation detailing the three Pulumi patterns used throughout the repository.

## Experiments and Documentation

### experiments/

This directory provides concrete usage patterns through minimal, single-command examples:

- **`experiments/tutorials/`** – Contains [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) and [`hello_world.py`](https://github.com/marin-community/marin/blob/main/hello_world.py) for new contributor onboarding.
- **`experiments/tokenize/`** – Scripts to test and validate tokenization pipelines.
- **`experiments/sft/`** – Supervised fine-tuning utilities and example pipelines.

### docs/

All user-facing documentation sources are built by MkDocs using [`mkdocs.yml`](https://github.com/marin-community/marin/blob/main/mkdocs.yml):

- **`docs/tutorials/`** – Step-by-step guides including [`docs/tutorials/installation.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/installation.md) for environment setup.
- **`docs/system-prompts/`** – Prompt templates used by the platform's AI agents.
- **`docs/reports/`** – Benchmark tables and experiment result summaries.

## Configuration and Testing

The `config/` directory stores environment-specific defaults loaded lazily at runtime. The top-level `tests/` directory contains integration tests for VLLM interfaces, while each package in `lib/` maintains its own unit test suite. GitHub-specific automation resides in `.github/`, including workflow definitions that trigger TPU CI jobs defined in `infra/tpu-ci/`.

### Navigating Python Import Paths

The directory structure directly determines import semantics across the layered architecture:

```python

# Import from the high-level pipeline package

from marin.execution.step_runner import StepRunner

# Access JAX optimizer configurations from the training library

from levanter.optim import AdamConfig

# Use dataset readers from the I/O utilities layer

from zephyr.readers import ParquetReader

```

These imports demonstrate the strict dependency direction: `marin` imports from `levanter`, and `zephyr` provides utilities to both, but no lower layer imports from a higher one.

## Summary

- The marin repository directory structure separates concerns into `lib/` (core code), `infra/` (deployment), `experiments/` (examples), and `docs/` (guides).
- Six core packages in `lib/` form a strict dependency hierarchy from low-level services (rigging) to high-level pipelines (marin).
- Each library package contains `src/`, `tests/`, and [`AGENTS.md`](https://github.com/marin-community/marin/blob/main/AGENTS.md) for isolated development.
- Infrastructure is defined as code using Pulumi in `infra/`, with specific stacks for TPU profiling and CI management.
- Configuration files in `config/` and integration tests in `tests/` support runtime flexibility and validation.

## Frequently Asked Questions

### What is the purpose of the lib/ directory in the marin repository?

The `lib/` directory contains the six core Python packages that implement the entire training stack. Each subdirectory (levanter, marin, iris, zephyr, fray, rigging) operates as a separate sub-project with its own source, tests, and agent guidelines, enforcing architectural boundaries between training primitives, pipeline logic, and infrastructure orchestration.

### How does the marin repository enforce clean dependency directions?

The repository architecture prohibits lower-level packages from importing higher-level ones. For example, `lib/levanter/` provides optimizers and checkpoints used by `lib/marin/`, but never imports from it. Similarly, `lib/zephyr/` handles dataset I/O without awareness of the pipeline framework, ensuring that foundational utilities remain independent of specific experiment implementations.

### Where are the Pulumi infrastructure definitions located?

All infrastructure-as-code definitions reside in `infra/`, specifically within `infra/xprof/` for profiling services, `infra/tpu-ci/` for continuous integration pods, and `infra/ducky/` for validation examples. The file [`infra/pulumi.md`](https://github.com/marin-community/marin/blob/main/infra/pulumi.md) documents the three architectural patterns used across these stacks.

### How do I find example training scripts in the marin repository?

Runnable examples live in `experiments/tutorials/` and include [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) at [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py), which demonstrates end-to-end tokenization, training, and execution in a single minimal script. Additional specialized examples for tokenization and supervised fine-tuning are available in `experiments/tokenize/` and `experiments/sft/` respectively.