Core Dependencies for Marin: Understanding the Multi-Package Python Workspace
The Marin repository is built on eleven core workspace packages—including marin-iris for job orchestration, marin-levanter for distributed JAX training, and marin-ducky for SQL services—that are declared in the root pyproject.toml to form a complete machine learning execution engine.
The marin-community/marin project is architected as a multi-package Python workspace where functionality is partitioned into specialized internal libraries. Examining the core dependencies for Marin reveals a layered system spanning from low-level tensor operations to high-level experiment orchestration, all wired together through the root configuration and the StepRunner execution engine.
Core Workspace Packages in Marin
The foundation of the project consists of internally maintained workspace packages declared in pyproject.toml. These packages are not external PyPI libraries but rather sub-projects within the Marin repository itself.
Job Orchestration and Distributed Execution
Two packages handle the execution fabric:
-
marin-iris: Defined at lines 12‑13 of
pyproject.toml, this package manages job orchestration and cluster resource allocation. It provides the scheduling primitives that determine where and when experiments run. -
marin-fray: Listed at lines 13‑14, this serves as the distributed execution framework, handling the actual dispatch and coordination of tasks across compute nodes.
Machine Learning and Training Libraries
The ML stack relies on JAX-based components:
-
marin-levanter: Specified at lines 15‑16, this is the large-scale JAX training library that implements optimization loops and distributed training strategies. Production scripts typically import
levanter.optim.AdamConfigfrom this package. -
marin-haliax: Found at lines 14‑15, this provides tensor and array utilities that support the numerical operations required by the training frameworks.
Data Processing and Storage Services
Data handling is split across three specialized packages:
-
marin-zephyr: Declared at lines 22‑23, this implements the dataset processing pipeline, handling ingestion and transformation workflows.
-
marin-ducky: Located at lines 32‑34, this provides a DuckDB-based ad-hoc SQL service for querying datasets directly within the Marin ecosystem.
-
marin-dupekit: Defined across lines 36‑43, this contains text deduplication utilities for cleaning training corpora.
Infrastructure and Core Utilities
Supporting infrastructure is provided by:
-
marin-core: Listed at lines 16‑17, this contains core data structures and shared utilities used across all other workspace packages.
-
marin-rigging: Found at lines 20‑22 with the
[secrets]extra, this handles infrastructure rigging and optional secret management for secure deployments. -
marin-finelog: Specified at lines 23‑24, this provides the logging client and CLI tooling for experiment monitoring.
-
watchdog: A lightweight but required top-level dependency at lines 35‑36, this monitors file-system events to trigger workflow updates.
Dependency Declaration in pyproject.toml
All core dependencies are explicitly declared in the root pyproject.toml. Unlike typical Python projects that rely solely on external PyPI packages, Marin’s essential dependencies are workspace packages—internal codebases that are installed in editable mode during development. The file maps each package name to its local directory, allowing the multi-repo structure to function as a unified system.
For external integrations requiring specialized hardware, the repository also maintains lib/marin/src/marin/external_dependencies.py. This file defines Git-based dependencies (e.g., vLLM GPU wheels and TPU inference libraries) that complement the core packages without bloating the minimal installation.
Runtime Integration of Core Dependencies
The interplay between these packages is demonstrated in typical experiment scripts. The following example shows how StepRunner, train_lm, and Levanter optimizers work together:
# Example: Running a simple Marin experiment using core packages
from marin.execution.step_runner import StepRunner
from marin.execution.lazy import lower
from marin.experiment.train import train_lm
from levanter.optim import AdamConfig
from experiments.llama import llama_nano
from experiments.marin_tokenizer import marin_tokenizer
from marin.experiment.data import tokenized
# 1️⃣ Tokenize a small dataset (uses core `marin-experiment` utilities)
tinystories = tokenized(
name="tokenized/tinystories",
source="roneneldan/TinyStories",
tokenizer=marin_tokenizer,
sample_count=500,
)
# 2️⃣ Define a training step (relies on core `marin-levanter` optimizer)
model_step = train_lm(
name="checkpoints/marin-nano-tinystories",
version="v1",
model=llama_nano,
optimizer=AdamConfig(learning_rate=5e-4, weight_decay=0.1),
datasets={tinystories: 1.0},
batch_size=4,
seq_len=2048,
num_train_steps=50,
resources="cpu", # works with core `marin-iris` resource configs
)
# 3️⃣ Execute the pipeline
if __name__ == "__main__":
StepRunner().run([lower(model_step)])
This script demonstrates the dependency graph: marin.execution.step_runner (from lib/marin/src/marin/execution/step_runner.py) orchestrates the workflow, marin.experiment provides the training DSL, and levanter.optim supplies the mathematical primitives—all enabled by the core packages declared in pyproject.toml.
External Dependencies and Optional Features
Beyond the eleven core workspace packages, the pyproject.toml contains an override-dependencies section that pins third-party libraries like datasets, equinox, ray, tiktoken, and rich. These are required for specific features—such as Hugging Face dataset integration or TPU-based inference—but are not part of the minimal execution engine.
The boundary between core and optional is intentional: a minimal installation of Marin requires only the workspace packages and watchdog, while full functionality demands the complete dependency tree defined in the override section.
Summary
- Marin is a multi-package workspace with eleven core internal dependencies declared in the root
pyproject.toml. - marin-iris handles cluster orchestration while marin-fray manages distributed execution.
- marin-levanter provides JAX-based training capabilities alongside marin-haliax for tensor operations.
- marin-ducky and marin-zephyr cover SQL services and data pipelines, respectively.
- StepRunner in
lib/marin/src/marin/execution/step_runner.pyserves as the central runtime that wires these dependencies together.
Frequently Asked Questions
What is the primary role of marin-iris in the core dependencies?
marin-iris serves as the job orchestration and cluster management layer, declared at lines 12‑13 of pyproject.toml. It determines resource allocation and scheduling policies for distributed experiments, working closely with marin-fray to dispatch tasks to available compute nodes according to the hardware requirements specified in experiment configurations.
How does marin-levanter differ from marin-fray?
While marin-fray provides the general-purpose distributed execution framework for task scheduling and cluster communication, marin-levanter (lines 15‑16) is specifically optimized for large-scale machine learning, implementing JAX-based training loops, checkpointing, and optimization strategies. In the execution flow, Fray moves the computation graph while Levanter performs the mathematical operations on accelerators.
Where are the core dependencies declared in the Marin repository?
All core workspace packages are declared in the [project.dependencies] section of the root [pyproject.toml](https://github.com/marin-community/marin/blob/main/pyproject.toml), with specific line references for each package (e.g., lines 12‑24 for the primary stack). Optional external dependencies are listed in the override-dependencies section of the same file, while Git-based hardware-specific dependencies are defined in lib/marin/src/marin/external_dependencies.py.
Is watchdog a required dependency for basic Marin functionality?
Yes, watchdog is a lightweight but required top-level dependency listed at lines 35‑36 of pyproject.toml. It provides file-system event monitoring that triggers workflow updates and experiment refreshes, and it is the only non-workspace package included in the minimal core dependency set required to run the Marin execution engine.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →