What Are the Key Libraries in Marin? Inside the Python Workspace Architecture
The Marin platform is composed of thirteen specialized Python libraries—including marin-iris for cluster orchestration, marin-levanter for JAX-based LLM training, and marin-haliax for typed tensor operations—that are organized as a workspace under the lib/ directory and registered in the top-level pyproject.toml.
The marin-community/marin repository structures its core functionality as a tightly-coupled Python workspace rather than a monolithic application. Each library lives in its own subdirectory under lib/ and is explicitly declared as a workspace member in pyproject.toml, providing discrete building blocks for distributed computing, machine learning, and data infrastructure.
Understanding the Marin Workspace Structure
Unlike traditional monorepos, Marin uses a native Python workspace defined in the root pyproject.toml. This file lists every library under the [tool.hatch.workspace] section, enabling shared dependency resolution while maintaining clean package boundaries. Source code for each library resides in lib/<library-name>/src/<library-name>/, following the src layout convention. For example, the orchestration layer marin-iris exposes its public API through lib/iris/src/iris/__init__.py.
Core Orchestration and Distributed Execution
These libraries manage job scheduling, cluster lifecycles, and distributed pipeline execution across compute resources.
marin-iris
marin-iris provides job orchestration, cluster management, and RPC services. It handles the creation and teardown of compute clusters, exposing a high-level Python API that abstracts underlying container orchestration.
from iris import cluster
# Start a temporary local cluster (uses Docker under the hood)
c = cluster.Cluster(name="demo")
c.up() # → creates pods / services
print(c.status()) # → shows health of the cluster
c.down() # → clean shutdown
The public API is defined in lib/iris/src/iris/, with cluster lifecycle methods implemented in the Cluster class.
marin-fray
marin-fray serves as the distributed execution engine that runs pipelines on clusters. It integrates with marin-iris to schedule tasks across nodes, managing the actual compute graph execution after resources are provisioned.
Machine Learning Training and Tensor Utilities
This category contains the JAX-based training framework and the typed array library that underpin Marin's ML capabilities.
marin-levanter
marin-levanter is a JAX-based training framework designed for large language models. It provides the Trainer class and configuration utilities that handle distributed training loops, checkpointing, and device mesh management.
from levanter import Trainer, ModelConfig
cfg = ModelConfig(
vocab_size=50257,
hidden_dim=256,
num_layers=4,
)
trainer = Trainer(config=cfg, data_path="my_dataset")
trainer.train(steps=500) # runs a JAX-accelerated training loop
The core training primitives reside in lib/levanter/src/levanter/trainer.py, which implements the gradient update loop and mixed-precision training logic.
marin-haliax
marin-haliax supplies typed tensor and array utilities used across the entire stack. It introduces named axes via the Axis class and the Array abstraction, preventing shape errors through runtime axis checking rather than implicit positional indices.
import haliax as hx
import jax.numpy as jnp
# Declare a named axis
Batch = hx.Axis("batch", 8)
# Create a tensor with that axis attached
x = hx.Array(jnp.arange(8), dims=(Batch,))
print(x.shape) # → (batch=8,)
The Array implementation is located in lib/haliax/src/haliax/array.py, providing the foundation for shape-safe neural network operations.
Data Processing and Storage Infrastructure
These libraries handle dataset streaming, persistent metadata storage, SQL querying, and text deduplication.
marin-zephyr
marin-zephyr manages dataset processing, sharding, and streaming utilities. It handles the ingestion of training data, implementing efficient sharding strategies for distributed training jobs. Dataset handling utilities are implemented in lib/zephyr/src/zephyr/dataset.py.
marin-finestore
marin-finestore acts as the persistent storage layer for metadata and experiment artifacts. It provides abstractions for saving and retrieving model checkpoints, configuration histories, and evaluation metrics across training runs.
marin-dupekit
marin-dupekit is a pure-Python text deduplication library with an optional Rust backend. It optimizes preprocessing pipelines by identifying and removing duplicate documents from large-scale text corpora before training.
marin-ducky
marin-ducky provides a DuckDB-backed ad-hoc SQL service used for dashboards and data runners. It enables efficient analytical queries over experiment metadata and training logs without requiring a separate database cluster, with the service implementation located in lib/ducky/src/ducky/sql_service.py.
Observability, UI, and System Glue
These packages provide logging infrastructure, web interfaces, and core shared abstractions.
marin-finelog
marin-finelog consists of a structured logging client (pure-Python) and a deployable log server. It standardizes telemetry across all Marin services, with the client implementation available in lib/finelog/src/finelog/logger.py.
marin-marina
marin-marina functions as the multi-app web kernel that serves the platform’s UI. Located in infra/marina/src/marina/app.py, it acts as the glue layer between the backend libraries and the frontend, routing requests to appropriate services.
marin-core
marin-core defines core abstractions including configuration management, metrics collection, and common types used throughout the workspace. It serves as the foundational dependency for nearly all other Marin libraries.
marin-rigging
marin-rigging provides helper utilities for deployment, secrets handling, and resource wiring. It simplifies the process of connecting Marin services to external infrastructure such as cloud storage buckets and credential stores.
Developer and Deployment Tools
marin-deploy
marin-deploy supplies CLI tools for building and publishing wheels, managing containers, and handling release logistics. It automates the packaging workflow for the workspace, ensuring consistent artifact generation across the library ecosystem.
Summary
- marin-iris, marin-fray, and marin-core provide the foundation for cluster orchestration and distributed execution.
- marin-levanter and marin-haliax enable JAX-based LLM training with typed tensor abstractions located in
lib/levanter/src/levanter/trainer.pyandlib/haliax/src/haliax/array.py. - marin-zephyr, marin-finestore, marin-dupekit, and marin-ducky handle the data lifecycle from ingestion and deduplication to SQL querying via DuckDB.
- marin-finelog, marin-marina, and marin-rigging provide observability, web UI serving, and deployment utilities.
- All thirteen libraries are registered as workspace members in the root
pyproject.tomland follow thelib/<name>/src/<name>/source layout.
Frequently Asked Questions
How is the Marin workspace structured in the repository?
The Marin repository uses a Python workspace defined in the root pyproject.toml, where each library is listed under workspace members and resides in its own directory under lib/. This structure allows shared dependency management while maintaining strict package boundaries, with source code following the src/ layout convention (e.g., lib/iris/src/iris/).
Which library handles LLM training in Marin?
marin-levanter is the dedicated library for large language model training, built on JAX. It provides the Trainer class and ModelConfig utilities for defining model architectures and executing distributed training loops across GPU or TPU clusters.
What distinguishes marin-haliax from standard JAX arrays?
marin-haliax introduces named axes via the Axis class and the Array abstraction, allowing tensors to carry semantic axis names (e.g., batch=8) rather than relying solely on positional indices. This prevents shape-related bugs by enforcing explicit axis alignment during operations, with the core implementation in lib/haliax/src/haliax/array.py.
Where does Marin store experiment metadata and artifacts?
marin-finestore serves as the persistent storage layer for metadata and experiment artifacts, while marin-ducky provides DuckDB-backed SQL services for querying this data. Finestore handles the write path for checkpoints and metrics, whereas Ducky enables analytical dashboards and ad-hoc queries over the stored metadata.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →