Marin Project Structure: A Deep Dive into the Monorepo Architecture
The marin project is organized as a Python-centric monorepo containing tightly-coupled libraries (lib/), experiment scripts (experiments/), infrastructure definitions (infra/), and comprehensive documentation (docs/), with core data processing handled by the Zephyr lazy dataset engine and distributed utilities provided by the Rigging library.
The marin-community/marin repository follows a layered architecture designed for large-scale machine learning pipelines. Understanding the marin project structure is essential for contributors working with its declarative data processing engine, distributed execution primitives, or pipeline orchestration tools.
Top-Level Directory Layout
The repository root divides functionality into seven primary directories, each serving a distinct purpose in the ML lifecycle.
lib/– Reusable Python libraries powering Marin’s pipeline, dataset handling, and distributed execution.experiments/– Example scripts demonstrating model training, tokenization, and scaling-law sweeps.docs/– User-facing knowledge base including installation guides and design overviews.infra/– Infrastructure-as-code (Pulumi) and ancillary services like the XProf profiling server.tests/– Automated verification split by sub-project (e.g.,tests/vllm/,tests/transform/).config/– Default YAML configuration files driving pipelines and external tools.
Key repository metadata files sit at the root: AGENTS.md defines global conventions, README.md provides the high-level overview, and CONTRIBUTING.md outlines the contribution workflow.
Core Libraries in lib/
The lib/ directory contains the foundational components of the marin project structure, implementing a three-layer architecture where Rigging provides low-level primitives, Zephyr builds the dataset engine on top, and Marin orchestrates the pipeline.
Zephyr: Lazy Dataset Processing
Located in lib/zephyr/, Zephyr implements a declarative, actor-based pipeline for large-scale dataset manipulation. According to lib/zephyr/AGENTS.md, all data flow between stages is filesystem-backed, with each stage writing output to a PickleDiskChunk on disk to minimize memory pressure.
Core implementation files include:
src/zephyr/dataset.py– Defines theDatasetclass and transformation primitives (group_by,deduplicate, etc.).src/zephyr/context.py– ManagesZephyrContextand the lifecycle of worker pools.src/zephyr/coordinator.py– Orchestrates task scheduling, counters, and result aggregation.src/zephyr/worker.py– Implements persistent workers that pull tasks, heartbeat, and report results.src/zephyr/plan.py– Generates physical execution plans and performs operation fusion.src/zephyr/readers.py&src/zephyr/writers.py– IO adapters for JSONL, Parquet, and Vortex formats.src/zephyr/shuffle.py– Scatter-merge shuffle implementation used across pipelines.
Workers communicate with the coordinator via RPCs that must block on .result() calls to preserve ordering and avoid race conditions.
Rigging: Distributed System Utilities
The lib/rigging/ library supplies reusable building blocks for distributed execution and observability, validated by tests in lib/rigging/tests/ such as test_distributed_lock.py.
Key submodules include:
rigging/tunnel.py– Secure RPC tunnelling between processes.rigging/token_authority.py– Handles secure token issuance for service-to-service authentication.rigging/telemetry/– Prometheus exporters and custom probes (NVIDIA, NCCL).rigging/testing.py– Test helpers for simulated clusters.rigging/server_auth.py– Server-side authentication middleware.
Marin: Pipeline Orchestration
While the top-level marin package source is largely vendored in lib/marin/, the directory enforces strict conventions documented in lib/marin/AGENTS.md:
- Use
fsspec.openfor all filesystem access to maintain cloud agnosticism. - Stream large GCS objects rather than copying them locally.
- Reference pipeline steps instead of hard-coding bucket paths.
Experiments and Tutorials
The experiments/ directory contains end-to-end examples demonstrating typical workflows. The file experiments/tutorials/train_tiny_model.py showcases tokenizing a small dataset, building a StepRunner, and launching a training step using Zephyr for tokenization and Marin for step definition.
Additional scripts handle scaling-law sweeps that drive large-scale runs on TPU pods, leveraging the same pipeline abstractions found in the core libraries.
Infrastructure and Configuration
Infrastructure Definitions
The infra/ directory houses Pulumi configurations for auxiliary services. The XProf profiling server at infra/xprof/server.py enables reproducible deployment of services required during large-scale training runs.
Configuration Files
Default pipeline configurations reside in config/marin.yaml, which drives experiments and external tools without requiring hard-coded parameters in the source.
Testing Strategy
The marin project structure splits testing across domain-specific directories to ensure comprehensive coverage of distributed components:
tests/vllm/– Integration tests for the VLLM inference backend, includingtests/vllm/test_llm_inference.py.tests/transform/– Unit tests for data-conversion utilities (HTML, HuggingFace datasets).- Library-specific tests – Co-located with source code (e.g.,
lib/zephyr/tests/,lib/rigging/tests/).
Execute the full test suite using the project-wide pre-commit wrapper:
./infra/pre-commit.py --all-files --fix
uv run --no-project infra/ci/run_tests.py
Documentation Organization
User-facing documentation lives under docs/ with tutorials organized by complexity:
docs/tutorials/installation.md– Environment setup and dependency installation.docs/tutorials/first-experiment.md– Step-by-step guide mirroringtrain_tiny_model.py.docs/tutorials/train-an-lm.md– Documentation for the "Delphi" scaling recipe.
Summary
- The marin project structure uses a monorepo layout with clear separation between libraries (
lib/), experiments (experiments/), and infrastructure (infra/). - Zephyr provides the lazy dataset processing engine with filesystem-backed dataflow, while Rigging supplies distributed primitives like tunneling and telemetry.
- Pipeline orchestration follows strict conventions using
fsspec.openand step-based configuration rather than hard-coded paths. - Testing is partitioned by concern (VLLM, transforms) alongside library-specific unit tests.
- Documentation and configuration are first-class citizens, with YAML configs in
config/and tutorials indocs/tutorials/.
Frequently Asked Questions
What is the purpose of the lib/ directory in the marin repository?
The lib/ directory contains the reusable Python libraries that constitute the core of the marin project structure. It houses Zephyr for lazy dataset processing, Rigging for distributed system utilities, and the vendored Marin orchestration layer. This centralized approach allows tight integration between the dataset engine and distributed execution primitives while maintaining modular boundaries.
How does Zephyr handle memory management during large-scale data processing?
According to lib/zephyr/AGENTS.md, Zephyr uses a filesystem-backed architecture where each pipeline stage writes its output to a PickleDiskChunk on disk rather than keeping data in memory. The next stage reads from these chunks, minimizing in-memory pressure across the declarative pipeline. Workers coordinate via RPCs that block on .result() calls to maintain execution ordering without consuming excessive RAM.
Where are the end-to-end training examples located?
Complete training examples reside in experiments/tutorials/, specifically experiments/tutorials/train_tiny_model.py. This script demonstrates tokenizing a dataset with Zephyr, constructing a StepRunner, and executing a training step. The docs/tutorials/ directory contains complementary written guides such as first-experiment.md that walk through the same workflow step-by-step.
How is authentication handled between distributed services in marin?
Authentication is managed by the Rigging library in lib/rigging/. The rigging/token_authority.py module handles secure token issuance for service-to-service authentication, while rigging/server_auth.py provides server-side authentication middleware. These components are tested in lib/rigging/tests/test_distributed_lock.py and related files to ensure secure communication across distributed training clusters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →