How to Build Marin from Source: Complete Installation Guide
To build Marin from source, clone the repository, create a Python 3.12+ virtual environment using uv, install dependencies with uv sync --all-packages and the appropriate hardware extra (cpu, gpu, or tpu), optionally compile Rust extensions via make rust-dev, and configure the required environment variables before running the tutorial script.
Marin is a modular research platform designed for training large language models at scale. Its architecture comprises several specialized layers—from the JAX-based Levanter training library to the Iris job orchestration system and Zephyr data processing layer. Building Marin from source allows you to modify these core components and run experiments on CPUs, GPUs, or TPUs according to the procedures documented in docs/tutorials/installation.md.
Prerequisites for Building Marin
Before compiling, ensure your system meets the baseline requirements. Marin requires Python 3.12 or newer, the uv package manager for dependency resolution, and Git for source control. While optional, a Rust toolchain is necessary if you intend to build native extensions from source rather than using pre-compiled wheels.
Step-by-Step Build Instructions
1. Clone the Repository
Fetch the latest source code from the marin-community organization:
git clone https://github.com/marin-community/marin.git
cd marin
2. Create a Python 3.12 Virtual Environment
Marin development relies on uv for environment management. Create and activate an isolated environment:
uv venv --python 3.12
source .venv/bin/activate
# Windows users: .venv\Scripts\activate
3. Install Core Dependencies with Hardware Extras
Install the full package stack in editable mode. You must specify the hardware target to pull the correct JAX wheels as declared in pyproject.toml:
# CPU-only installation
uv sync --all-packages
# For GPU support
uv sync --all-packages --extra=gpu
# For TPU support
uv sync --all-packages --extra=tpu
4. Build Rust Extensions from Source (Optional)
If you need to modify native code, switch from pre-built wheels to local Cargo builds using the Makefile targets:
make rust-dev
uv sync --all-packages
The make rust-dev target modifies pyproject.toml to reference local Cargo builds. To revert to standard wheels after development, run make rust-user.
5. Configure Environment Variables
Marin requires three critical environment variables for operation. Set these before running experiments:
export WANDB_API_KEY=your_key
export HF_TOKEN=your_hf_token
export MARIN_PREFIX=$HOME/marin_artifacts
The MARIN_PREFIX variable specifies an fsspec-compatible path for artifact storage, while WANDB_API_KEY and HF_TOKEN enable experiment tracking and gated model access via the Hugging Face Hub.
6. Verify the Build with the Tutorial
Run the tiny-model tutorial to confirm your installation succeeds:
wandb offline
uv run python experiments/tutorials/train_tiny_model.py \
--device cpu \
--dataset tinystories \
--version dev \
--run
This executes the minimal pipeline defined in experiments/tutorials/train_tiny_model.py, validating that Levanter, Iris, Zephyr, and the Marin orchestration layer are functioning correctly.
Understanding Marin’s Modular Architecture
Building from source requires awareness of how Marin’s components interact. The repository organizes functionality into distinct layers documented in their respective README files:
-
lib/levanter/– Contains the JAX-based training library providing model components, optimizers, and data pipelines. Referencelib/levanter/README.mdfor low-level APIs. -
lib/iris/– Implements job orchestration that abstracts cluster resources and manages automatic retries across CPUs, GPUs, and TPUs. Seelib/iris/README.mdfor scheduling documentation. -
lib/zephyr/– Provides dataset processing utilities including readers, writers, and the "vortex" in-memory abstraction for streaming large datasets. Consultlib/zephyr/README.mdfor data pipeline configurations. -
lib/marin/– The top-level pipeline framework that composes training steps as a DAG and manages lazy artifact materialization, integrating with Weights & Biases for experiment tracking.
Key Configuration Files
Several repository files govern the build process and validation:
-
docs/tutorials/installation.md– The authoritative source for installation procedures and troubleshooting. -
pyproject.toml– Declares Python dependencies, optional JAX extras for hardware acceleration, and Rust extension metadata. -
Makefile– Provides therust-devandrust-usertargets for switching between source and wheel-based native extensions. -
experiments/tutorials/train_tiny_model.py– The canonical verification script that exercises the full stack. -
tests/vllm/test_llm_inference.py– Example unit test for validating VLLM integration. -
.github/workflows/marin-lint.yaml– CI workflow demonstrating the exact steps used to validate successful builds in continuous integration.
Summary
Building Marin from source establishes a development environment matching the core maintainers' configuration. Key steps include:
- Cloning the repository and initializing a Python 3.12 virtual environment with
uv - Installing dependencies via
uv syncwith hardware-specific extras (--extra=cpu,gpu, ortpu) - Optionally compiling Rust extensions using
make rust-devwhen modifying native code - Configuring
WANDB_API_KEY,HF_TOKEN, andMARIN_PREFIXenvironment variables - Validating the build by executing
experiments/tutorials/train_tiny_model.py
Frequently Asked Questions
What Python version is required to build Marin?
Marin requires Python 3.12 or newer. Earlier versions are not supported due to dependency constraints in the JAX ecosystem and type annotation features used throughout lib/levanter/ and lib/marin/.
Do I need to install Rust to build Marin from source?
No, Rust is only required if you intend to modify or debug the native extensions. By default, uv sync installs pre-compiled wheels. Use make rust-dev only when you need to compile the Rust crates from source using Cargo, as configured in pyproject.toml.
How do I switch between CPU and GPU builds after initial installation?
Re-run uv sync --all-packages with the appropriate extra flag. For example, switch to GPU by executing uv sync --all-packages --extra=gpu, which updates the JAX installation in your virtual environment to use CUDA-enabled wheels defined in pyproject.toml.
Where does Marin store built artifacts and model checkpoints?
The MARIN_PREFIX environment variable controls artifact storage. Set this to any fsspec-compatible path (local filesystem or cloud storage URI). By default, the tutorial suggests $HOME/marin_artifacts, but production deployments typically use distributed storage paths for scalability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →