# Core Dependencies for Marin: Understanding the Multi-Package Python Workspace

> Discover the eleven core dependencies for Marin, the multi-package Python workspace that powers your machine learning execution engine. Learn about key packages like iris, levanter, and ducky.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-08-29

---

**The Marin repository is built on eleven core workspace packages—including marin-iris for job orchestration, marin-levanter for distributed JAX training, and marin-ducky for SQL services—that are declared in the root [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) to form a complete machine learning execution engine.**

The marin-community/marin project is architected as a multi-package Python workspace where functionality is partitioned into specialized internal libraries. Examining the core dependencies for Marin reveals a layered system spanning from low-level tensor operations to high-level experiment orchestration, all wired together through the root configuration and the `StepRunner` execution engine.

## Core Workspace Packages in Marin

The foundation of the project consists of internally maintained workspace packages declared in [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml). These packages are not external PyPI libraries but rather sub-projects within the Marin repository itself.

### Job Orchestration and Distributed Execution

Two packages handle the execution fabric:

- **marin-iris**: Defined at lines 12‑13 of [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml), this package manages job orchestration and cluster resource allocation. It provides the scheduling primitives that determine where and when experiments run.

- **marin-fray**: Listed at lines 13‑14, this serves as the distributed execution framework, handling the actual dispatch and coordination of tasks across compute nodes.

### Machine Learning and Training Libraries

The ML stack relies on JAX-based components:

- **marin-levanter**: Specified at lines 15‑16, this is the large-scale JAX training library that implements optimization loops and distributed training strategies. Production scripts typically import `levanter.optim.AdamConfig` from this package.

- **marin-haliax**: Found at lines 14‑15, this provides tensor and array utilities that support the numerical operations required by the training frameworks.

### Data Processing and Storage Services

Data handling is split across three specialized packages:

- **marin-zephyr**: Declared at lines 22‑23, this implements the dataset processing pipeline, handling ingestion and transformation workflows.

- **marin-ducky**: Located at lines 32‑34, this provides a DuckDB-based ad-hoc SQL service for querying datasets directly within the Marin ecosystem.

- **marin-dupekit**: Defined across lines 36‑43, this contains text deduplication utilities for cleaning training corpora.

### Infrastructure and Core Utilities

Supporting infrastructure is provided by:

- **marin-core**: Listed at lines 16‑17, this contains core data structures and shared utilities used across all other workspace packages.

- **marin-rigging**: Found at lines 20‑22 with the `[secrets]` extra, this handles infrastructure rigging and optional secret management for secure deployments.

- **marin-finelog**: Specified at lines 23‑24, this provides the logging client and CLI tooling for experiment monitoring.

- **watchdog**: A lightweight but required top-level dependency at lines 35‑36, this monitors file-system events to trigger workflow updates.

## Dependency Declaration in [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml)

All core dependencies are explicitly declared in the root [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml). Unlike typical Python projects that rely solely on external PyPI packages, Marin’s essential dependencies are **workspace packages**—internal codebases that are installed in editable mode during development. The file maps each package name to its local directory, allowing the multi-repo structure to function as a unified system.

For external integrations requiring specialized hardware, the repository also maintains [`lib/marin/src/marin/external_dependencies.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/external_dependencies.py). This file defines Git-based dependencies (e.g., vLLM GPU wheels and TPU inference libraries) that complement the core packages without bloating the minimal installation.

## Runtime Integration of Core Dependencies

The interplay between these packages is demonstrated in typical experiment scripts. The following example shows how `StepRunner`, `train_lm`, and Levanter optimizers work together:

```python

# Example: Running a simple Marin experiment using core packages

from marin.execution.step_runner import StepRunner
from marin.execution.lazy import lower
from marin.experiment.train import train_lm
from levanter.optim import AdamConfig
from experiments.llama import llama_nano
from experiments.marin_tokenizer import marin_tokenizer
from marin.experiment.data import tokenized

# 1️⃣ Tokenize a small dataset (uses core `marin-experiment` utilities)

tinystories = tokenized(
    name="tokenized/tinystories",
    source="roneneldan/TinyStories",
    tokenizer=marin_tokenizer,
    sample_count=500,
)

# 2️⃣ Define a training step (relies on core `marin-levanter` optimizer)

model_step = train_lm(
    name="checkpoints/marin-nano-tinystories",
    version="v1",
    model=llama_nano,
    optimizer=AdamConfig(learning_rate=5e-4, weight_decay=0.1),
    datasets={tinystories: 1.0},
    batch_size=4,
    seq_len=2048,
    num_train_steps=50,
    resources="cpu",  # works with core `marin-iris` resource configs

)

# 3️⃣ Execute the pipeline

if __name__ == "__main__":
    StepRunner().run([lower(model_step)])

```

This script demonstrates the dependency graph: `marin.execution.step_runner` (from [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py)) orchestrates the workflow, `marin.experiment` provides the training DSL, and `levanter.optim` supplies the mathematical primitives—all enabled by the core packages declared in [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).

## External Dependencies and Optional Features

Beyond the eleven core workspace packages, the [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) contains an *override-dependencies* section that pins third-party libraries like `datasets`, `equinox`, `ray`, `tiktoken`, and `rich`. These are required for specific features—such as Hugging Face dataset integration or TPU-based inference—but are not part of the minimal execution engine.

The boundary between core and optional is intentional: a minimal installation of Marin requires only the workspace packages and `watchdog`, while full functionality demands the complete dependency tree defined in the override section.

## Summary

- **Marin is a multi-package workspace** with eleven core internal dependencies declared in the root [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).
- **marin-iris** handles cluster orchestration while **marin-fray** manages distributed execution.
- **marin-levanter** provides JAX-based training capabilities alongside **marin-haliax** for tensor operations.
- **marin-ducky** and **marin-zephyr** cover SQL services and data pipelines, respectively.
- **StepRunner** in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py) serves as the central runtime that wires these dependencies together.

## Frequently Asked Questions

### What is the primary role of marin-iris in the core dependencies?

**marin-iris** serves as the job orchestration and cluster management layer, declared at lines 12‑13 of [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml). It determines resource allocation and scheduling policies for distributed experiments, working closely with `marin-fray` to dispatch tasks to available compute nodes according to the hardware requirements specified in experiment configurations.

### How does marin-levanter differ from marin-fray?

While **marin-fray** provides the general-purpose distributed execution framework for task scheduling and cluster communication, **marin-levanter** (lines 15‑16) is specifically optimized for large-scale machine learning, implementing JAX-based training loops, checkpointing, and optimization strategies. In the execution flow, Fray moves the computation graph while Levanter performs the mathematical operations on accelerators.

### Where are the core dependencies declared in the Marin repository?

All core workspace packages are declared in the `[project.dependencies]` section of the root [[`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml)](https://github.com/marin-community/marin/blob/main/pyproject.toml), with specific line references for each package (e.g., lines 12‑24 for the primary stack). Optional external dependencies are listed in the *override-dependencies* section of the same file, while Git-based hardware-specific dependencies are defined in [`lib/marin/src/marin/external_dependencies.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/external_dependencies.py).

### Is watchdog a required dependency for basic Marin functionality?

Yes, **watchdog** is a lightweight but required top-level dependency listed at lines 35‑36 of [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml). It provides file-system event monitoring that triggers workflow updates and experiment refreshes, and it is the only non-workspace package included in the minimal core dependency set required to run the Marin execution engine.