Dependency Hierarchy Between Iris, Levanter, Zephyr, and Marin: Runtime Stack Explained
The marin-community/marin repository organizes its machine learning stack into four distinct layers where Iris provides the distributed runtime foundation, Levanter and Zephyr build upon it for training and data processing respectively, and Marin orchestrates pipelines across all three lower-level libraries.
The dependency hierarchy between Iris, Levanter, Zephyr, and Marin defines how distributed jobs, model training, and data processing interact within the Marin ecosystem. Understanding this stack is essential for debugging import errors, optimizing pipeline performance, and extending functionality. This article traces the actual import relationships found in the source code to map the precise architectural boundaries.
The Layered Architecture
The codebase follows a strict layered architecture where each library occupies a specific position in the stack and only imports from layers beneath it.
-
Iris (Foundation Layer) – Located in
lib/iris/src/iris/, this library provides the distributed-job orchestration runtime. It exposesiris_ctxandget_iris_ctxfromlib/iris/src/iris/client/client.pyand maintains no dependencies on Levanter, Zephyr, or Marin, making it the pure base of the hierarchy. -
Levanter (Training Layer) – Positioned in
lib/levanter/src/levanter/, this JAX-based language-model training framework imports exclusively from Iris. Inlib/levanter/src/levanter/distributed.py, Levanter callsfrom iris.client.client import iris_ctxto register workers and report distributed status, while optionally utilizingiris.runtime.jax_initfor JAX-specific initialization. -
Zephyr (Data Processing Layer) – Found in
lib/zephyr/src/zephyr/, this lazy dataset-processing library primarily depends on Iris viafrom iris.client.client import get_iris_ctxinlib/zephyr/src/zephyr/context.py. While its core functionality only requires Iris, Zephyr optionally imports megascale utilities from Levanter (such aslevanter.megascale.configure_megascale_from_iris) when handling large-scale sharding strategies. -
Marin (Orchestration Layer) – As the top-level pipeline controller in
lib/marin/src/marin/, this library imports from all three underlying layers. The inference controller inlib/marin/src/marin/inference/iris.pydemonstrates this by consumingiris_ctxfrom Iris,train_lmfrom Levanter, andget_iris_ctxfrom Zephyr to coordinate end-to-end workflows.
How the Hierarchy Works
The dependency flow follows a strict bottom-up pattern that ensures lower-level utilities remain agnostic of higher-level domain logic.
-
Iris initializes the runtime. The
IrisClientclass and context managers inlib/iris/src/iris/client/client.pyprovide the telemetry and job-tracking APIs that all other layers consume without circular references. -
Levanter registers training jobs. When executing distributed training, Levanter’s
distributed.pymodule invokesiris_ctxto hook workers into the Iris runtime for progress reporting and cleanup signaling. -
Zephyr registers data tasks. Dataset processing steps utilize
get_iris_ctxfromlib/zephyr/src/zephyr/context.pyto integrate task-status bookkeeping with the same monitoring infrastructure used by training jobs. -
Marin orchestrates cross-library workflows. The pipeline controller imports functions from all three libraries to stitch together Iris job management, Levanter training steps, and Zephyr data-processing steps into unified execution graphs.
Code-Level Evidence
The following snippets demonstrate the exact import statements that enforce the hierarchy boundaries.
Iris exports the core context API:
# lib/iris/src/iris/client/client.py
from iris.client.client import iris_ctx, get_iris_ctx, IrisClient
Levanter consumes the Iris runtime:
# lib/levanter/src/levanter/distributed.py
from iris.client.client import iris_ctx
from iris.runtime.jax_init import initialize_iris_jax
Zephyr accesses Iris for task tracking:
# lib/zephyr/src/zephyr/context.py
from iris.client.client import get_iris_ctx
Marin imports from all three layers:
# lib/marin/src/marin/inference/iris.py
from iris.client.client import iris_ctx
from levanter.main.train_lm import train_lm
from zephyr.context import get_iris_ctx
Summary
- Iris sits at the bottom of the stack with zero dependencies on the other three libraries, providing
iris_ctxandget_iris_ctxfor distributed runtime management. - Levanter depends solely on Iris, importing
iris_ctxfromlib/levanter/src/levanter/distributed.pyto enable distributed JAX training. - Zephyr primarily depends on Iris via
get_iris_ctxinlib/zephyr/src/zephyr/context.py, with optional utility imports from Levanter for megascale operations. - Marin is the sole consumer of all three libraries, orchestrating pipelines that integrate Iris job control, Levanter training, and Zephyr data processing.
Frequently Asked Questions
Which library forms the base of the Marin dependency stack?
Iris forms the foundation. Located in lib/iris/src/iris/client/client.py, it exports iris_ctx and get_iris_ctx that both Levanter and Zephyr import for distributed runtime services, while maintaining no internal dependencies on the higher-level libraries.
Does Zephyr require Levanter to function?
No. Zephyr operates independently using only Iris for task-status bookkeeping via get_iris_ctx. However, when processing datasets at megascale, Zephyr may optionally import helper functions from Levanter such as configure_megascale_from_iris to optimize sharding strategies.
How does Marin differ from Iris in terms of dependencies?
While Iris imports from none of the other three libraries, Marin imports from all three. Specifically, lib/marin/src/marin/inference/iris.py combines iris_ctx (Iris), train_lm (Levanter), and get_iris_ctx (Zephyr) to build coordinated machine learning pipelines.
What specific Iris functions do Levanter and Zephyr consume?
Levanter consumes iris_ctx from lib/iris/src/iris/client/client.py primarily within lib/levanter/src/levanter/distributed.py for worker registration. Zephyr consumes get_iris_ctx from the same Iris client module within lib/zephyr/src/zephyr/context.py for task-level context management.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →