Main Modules and Packages in the Marin Repository: A Complete Architecture Guide

The Marin repository organizes its functionality into nine specialized Python libraries located under the top-level lib/ directory, forming a layered dependency stack that enables large-scale language model research from data ingestion through distributed training.

The marin-community/marin codebase is structured as a collection of loosely-coupled packages designed to handle specific stages of the ML lifecycle. Understanding the main modules or packages in Marin is essential for anyone contributing to or building upon this platform, as the architecture deliberately enforces unidirectional dependencies where higher-level libraries import only from lower-level ones.

Core Library Architecture and Design Philosophy

The Marin repository follows a strict hierarchical dependency model to maintain clean separation of concerns. The dependency flow runs from low-level data utilities upward to high-level orchestration: iris → levanter → zephyr → marin.

This architecture ensures that foundational utilities remain independent of business logic, while the top-level marin package can compose functionality from all underlying layers. All core libraries reside under the lib/ directory, with each package containing its source code in lib/<package>/src/<package>/.

The Nine Main Packages in Marin

Zephyr: Low-Level Data Pipeline Utilities

Zephyr provides the foundational data-processing layer, implementing readers, writers, workers, and I/O optimizations. Located at lib/zephyr, this package handles the initial data ingestion and transformation stages.

Key implementation files include lib/zephyr/src/zephyr/worker.py, which manages parallel data loading workers. Typical usage involves reading Parquet datasets directly from cloud storage:

from lib.zephyr import ParquetReader

reader = ParquetReader(path="gs://my-bucket/dataset.parquet")
for record in reader:
    # process each record …

    pass

Levanter: JAX-Based Training Framework

Levanter constitutes the computational engine of the stack, built on JAX for high-performance model training. Found at lib/levanter, it contains model definitions, trainer implementations, and optimizer configurations.

The primary entry point is lib/levanter/src/levanter/trainer.py, which defines the core training loops and checkpoint management. Levanter consumes data processed by Zephyr and provides the Trainer class used by higher-level orchestration.

Iris: Job Orchestration and Scheduling

Iris manages distributed execution across compute clusters, handling job scheduling and resource allocation. Located at lib/iris, this package bridges the gap between training logic and hardware infrastructure.

The scheduler implementation in lib/iris/src/iris/job_scheduler.py provides the JobScheduler class for submitting and monitoring training workloads on TPU or GPU clusters.

Marin: High-Level Pipeline Orchestration

The marin package itself (located at lib/marin) provides the highest-level API, gluing together Iris, Levanter, and Zephyr for end-to-end experiment management. It defines abstractions for experiments, configurations, and workflow composition.

The core abstraction resides in lib/marin/src/marin/experiment.py, which enables researchers to define complete training pipelines:

from lib.marin import Experiment, Config
from lib.levanter import Trainer
from lib.iris import JobScheduler

cfg = Config(
    model="gpt2-small",
    optimizer="adamw",
    dataset="openwebtext",
    batch_size=128,
)

trainer = Trainer(cfg)
scheduler = JobScheduler(cluster="tpu-v4")
experiment = Experiment(name="gpt2_small_demo", trainer=trainer)

scheduler.submit(experiment)

Fray: Distributed Execution Runtime

Fray implements the actor-model runtime for distributed computation, managing actors, futures, and client-server communication. Located at lib/fray, it provides the underlying execution primitives used by Iris for cluster management.

Key components are defined in lib/fray/src/fray/actor.py, which implements the actor-based concurrency model.

Rigging: Infrastructure and Utilities

Rigging supplies cross-cutting infrastructure concerns including authentication, secrets management, telemetry, and storage abstractions. Located at lib/rigging, it provides shared utilities used across the entire stack.

The secrets management system in lib/rigging/src/rigging/secrets.py exemplifies this utility:

from lib.rigging.secrets import SecretStore

store = SecretStore()
api_key = store.get("openai_api_key")

Ducky: RPC and Tunneling Framework

Ducky is a lightweight RPC and tunneling framework used primarily by Rigging and Fray for inter-service communication. Located at lib/ducky, it implements the networking layer required for distributed coordination.

Core tunneling logic resides in lib/ducky/src/ducky/tunnel.py.

Dupekit: Data Deduplication Utilities

Dupekit provides specialized algorithms for detecting and removing duplicate content in training datasets. Located at lib/dupekit, it implements hashing strategies including MinHash and Bloom filters for efficient deduplication at scale.

The MinHash implementation in lib/dupekit/src/dupekit/minhash.py serves as the primary interface for near-duplicate detection in data preprocessing pipelines.

Finestore: Persistent Artifact Storage

Finestore manages versioned storage for datasets and model artifacts, providing filesystem abstractions, caching layers, and schema management. Located at lib/finestore, it ensures reproducibility by tracking data lineage and model versions.

The storage interface is defined in lib/finestore/src/finestore/store.py.

Additional Repository Components

Beyond the core lib/ packages, the Marin repository includes several supporting directories:

  • experiments/ – Example training scripts and tutorials demonstrating package integration patterns
  • infra/ – Pulumi-based Infrastructure-as-Code definitions for deploying TPU clusters and GCP resources
  • docs/ – MkDocs-based documentation covering installation guides and architectural overviews

Summary

The Marin repository organizes its functionality into nine distinct Python libraries under lib/, each addressing specific requirements of large-scale ML research:

This layered architecture, with its strict dependency hierarchy from Zephyr up to Marin, enables scalable, reproducible language model research while maintaining clear separation of concerns between data processing, training computation, and orchestration logic.

Frequently Asked Questions

What is the purpose of the Levanter package in Marin?

Levanter serves as the JAX-based training framework within Marin, handling model definitions, optimization loops, and distributed training logic. It consumes preprocessed data from Zephyr and exposes trainer implementations that Iris schedules across compute clusters. The core training logic resides in lib/levanter/src/levanter/trainer.py.

How does the Marin package differ from the Iris package?

While Iris focuses specifically on job scheduling and cluster resource management, Marin provides high-level pipeline orchestration that coordinates multiple subsystems. Marin imports from both Iris (for scheduling) and Levanter (for training) to compose end-to-end experiments, whereas Iris operates independently at the infrastructure layer managing job lifecycle on TPU/GPU clusters.

Where are data deduplication utilities located in the Marin codebase?

Data deduplication functionality resides in the Dupekit package at lib/dupekit. This module implements MinHash algorithms and Bloom filters specifically designed for preprocessing large training datasets to remove duplicate content before ingestion into the training pipeline, with core algorithms in lib/dupekit/src/dupekit/minhash.py.

What file path contains the core Experiment class for defining training jobs?

The Experiment class is defined in lib/marin/src/marin/experiment.py. This class serves as the primary abstraction for configuring and launching complete training workflows, integrating configurations from Levanter with scheduling capabilities from Iris to provide the high-level API for the Marin platform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →