# What Are the Key Libraries in Marin? Inside the Python Workspace Architecture

> Explore the key Python libraries in Marin, including marin-iris, marin-levanter, and marin-haliax. Understand the Marin workspace architecture and its specialized components for efficient LLM training and cluster orchestration.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: deep-dive
- Published: 2026-09-10

---

**The Marin platform is composed of thirteen specialized Python libraries—including marin-iris for cluster orchestration, marin-levanter for JAX-based LLM training, and marin-haliax for typed tensor operations—that are organized as a workspace under the `lib/` directory and registered in the top-level [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).**

The `marin-community/marin` repository structures its core functionality as a tightly-coupled Python workspace rather than a monolithic application. Each library lives in its own subdirectory under **`lib/`** and is explicitly declared as a workspace member in **[`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml)**, providing discrete building blocks for distributed computing, machine learning, and data infrastructure.

## Understanding the Marin Workspace Structure

Unlike traditional monorepos, Marin uses a native Python workspace defined in the root **[`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml)**. This file lists every library under the `[tool.hatch.workspace]` section, enabling shared dependency resolution while maintaining clean package boundaries. Source code for each library resides in `lib/<library-name>/src/<library-name>/`, following the `src` layout convention. For example, the orchestration layer `marin-iris` exposes its public API through [`lib/iris/src/iris/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/iris/src/iris/__init__.py).

## Core Orchestration and Distributed Execution

These libraries manage job scheduling, cluster lifecycles, and distributed pipeline execution across compute resources.

### marin-iris

**marin-iris** provides job orchestration, cluster management, and RPC services. It handles the creation and teardown of compute clusters, exposing a high-level Python API that abstracts underlying container orchestration.

```python
from iris import cluster

# Start a temporary local cluster (uses Docker under the hood)

c = cluster.Cluster(name="demo")
c.up()                     # → creates pods / services

print(c.status())          # → shows health of the cluster

c.down()                   # → clean shutdown

```

The public API is defined in `lib/iris/src/iris/`, with cluster lifecycle methods implemented in the `Cluster` class.

### marin-fray

**marin-fray** serves as the distributed execution engine that runs pipelines on clusters. It integrates with marin-iris to schedule tasks across nodes, managing the actual compute graph execution after resources are provisioned.

## Machine Learning Training and Tensor Utilities

This category contains the JAX-based training framework and the typed array library that underpin Marin's ML capabilities.

### marin-levanter

**marin-levanter** is a JAX-based training framework designed for large language models. It provides the `Trainer` class and configuration utilities that handle distributed training loops, checkpointing, and device mesh management.

```python
from levanter import Trainer, ModelConfig

cfg = ModelConfig(
    vocab_size=50257,
    hidden_dim=256,
    num_layers=4,
)

trainer = Trainer(config=cfg, data_path="my_dataset")
trainer.train(steps=500)   # runs a JAX-accelerated training loop

```

The core training primitives reside in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py), which implements the gradient update loop and mixed-precision training logic.

### marin-haliax

**marin-haliax** supplies typed tensor and array utilities used across the entire stack. It introduces named axes via the `Axis` class and the `Array` abstraction, preventing shape errors through runtime axis checking rather than implicit positional indices.

```python
import haliax as hx
import jax.numpy as jnp

# Declare a named axis

Batch = hx.Axis("batch", 8)

# Create a tensor with that axis attached

x = hx.Array(jnp.arange(8), dims=(Batch,))
print(x.shape)   # → (batch=8,)

```

The `Array` implementation is located in [`lib/haliax/src/haliax/array.py`](https://github.com/marin-community/marin/blob/main/lib/haliax/src/haliax/array.py), providing the foundation for shape-safe neural network operations.

## Data Processing and Storage Infrastructure

These libraries handle dataset streaming, persistent metadata storage, SQL querying, and text deduplication.

### marin-zephyr

**marin-zephyr** manages dataset processing, sharding, and streaming utilities. It handles the ingestion of training data, implementing efficient sharding strategies for distributed training jobs. Dataset handling utilities are implemented in [`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py).

### marin-finestore

**marin-finestore** acts as the persistent storage layer for metadata and experiment artifacts. It provides abstractions for saving and retrieving model checkpoints, configuration histories, and evaluation metrics across training runs.

### marin-dupekit

**marin-dupekit** is a pure-Python text deduplication library with an optional Rust backend. It optimizes preprocessing pipelines by identifying and removing duplicate documents from large-scale text corpora before training.

### marin-ducky

**marin-ducky** provides a DuckDB-backed ad-hoc SQL service used for dashboards and data runners. It enables efficient analytical queries over experiment metadata and training logs without requiring a separate database cluster, with the service implementation located in [`lib/ducky/src/ducky/sql_service.py`](https://github.com/marin-community/marin/blob/main/lib/ducky/src/ducky/sql_service.py).

## Observability, UI, and System Glue

These packages provide logging infrastructure, web interfaces, and core shared abstractions.

### marin-finelog

**marin-finelog** consists of a structured logging client (pure-Python) and a deployable log server. It standardizes telemetry across all Marin services, with the client implementation available in [`lib/finelog/src/finelog/logger.py`](https://github.com/marin-community/marin/blob/main/lib/finelog/src/finelog/logger.py).

### marin-marina

**marin-marina** functions as the multi-app web kernel that serves the platform’s UI. Located in [`infra/marina/src/marina/app.py`](https://github.com/marin-community/marin/blob/main/infra/marina/src/marina/app.py), it acts as the glue layer between the backend libraries and the frontend, routing requests to appropriate services.

### marin-core

**marin-core** defines core abstractions including configuration management, metrics collection, and common types used throughout the workspace. It serves as the foundational dependency for nearly all other Marin libraries.

### marin-rigging

**marin-rigging** provides helper utilities for deployment, secrets handling, and resource wiring. It simplifies the process of connecting Marin services to external infrastructure such as cloud storage buckets and credential stores.

## Developer and Deployment Tools

### marin-deploy

**marin-deploy** supplies CLI tools for building and publishing wheels, managing containers, and handling release logistics. It automates the packaging workflow for the workspace, ensuring consistent artifact generation across the library ecosystem.

## Summary

- **marin-iris**, **marin-fray**, and **marin-core** provide the foundation for cluster orchestration and distributed execution.
- **marin-levanter** and **marin-haliax** enable JAX-based LLM training with typed tensor abstractions located in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py) and [`lib/haliax/src/haliax/array.py`](https://github.com/marin-community/marin/blob/main/lib/haliax/src/haliax/array.py).
- **marin-zephyr**, **marin-finestore**, **marin-dupekit**, and **marin-ducky** handle the data lifecycle from ingestion and deduplication to SQL querying via DuckDB.
- **marin-finelog**, **marin-marina**, and **marin-rigging** provide observability, web UI serving, and deployment utilities.
- All thirteen libraries are registered as workspace members in the root [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) and follow the `lib/<name>/src/<name>/` source layout.

## Frequently Asked Questions

### How is the Marin workspace structured in the repository?

The Marin repository uses a Python workspace defined in the root [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml), where each library is listed under workspace members and resides in its own directory under `lib/`. This structure allows shared dependency management while maintaining strict package boundaries, with source code following the `src/` layout convention (e.g., `lib/iris/src/iris/`).

### Which library handles LLM training in Marin?

**marin-levanter** is the dedicated library for large language model training, built on JAX. It provides the `Trainer` class and `ModelConfig` utilities for defining model architectures and executing distributed training loops across GPU or TPU clusters.

### What distinguishes marin-haliax from standard JAX arrays?

**marin-haliax** introduces named axes via the `Axis` class and the `Array` abstraction, allowing tensors to carry semantic axis names (e.g., `batch=8`) rather than relying solely on positional indices. This prevents shape-related bugs by enforcing explicit axis alignment during operations, with the core implementation in [`lib/haliax/src/haliax/array.py`](https://github.com/marin-community/marin/blob/main/lib/haliax/src/haliax/array.py).

### Where does Marin store experiment metadata and artifacts?

**marin-finestore** serves as the persistent storage layer for metadata and experiment artifacts, while **marin-ducky** provides DuckDB-backed SQL services for querying this data. Finestore handles the write path for checkpoints and metrics, whereas Ducky enables analytical dashboards and ad-hoc queries over the stored metadata.